UTHealth Houston-developed AI tool nears psychiatrist-level accuracy
Co-first authors Benson Mwangi Irungu, PhD, and Hammza Hamoudi, MD. (Photo by UTHealth Houston)
UTHealth Houston researchers developed an artificial intelligence system that can perform a mental health evaluation at nearly the same accuracy as a team of psychiatrists.
The results of their work, done in collaboration with Yale University and published this week in Nature’s npj Mental Health Research, brings the team a step closer to deploying the model in educational and clinical settings.
“The AI tool we built can perform a diagnostic evaluation on a patient that is almost as good as a team of psychiatrists. It’s very impressive,” said co-first author Hammza Hamoudi, MD, postdoctoral research fellow in the Department of Psychiatry and Behavioral Sciences at McGovern Medical School at UTHealth Houston.
The team combined the pretrained neural networks on a system called Qwen3-Omni with custom software they developed to analyze video recordings of patients’ speech, tone of voice, and behavior. The resulting system brings together observations across visits and generates written explanations for its mental-status assessments.
Senior author Cesar Soutullo, MD, PhD, vice chair and chief of Child and Adolescent Psychiatry in the Department of Psychiatry and Behavioral Sciences at McGovern Medical School, said the ultimate goal is to supplement, rather than replace, the work of psychiatrists.
“The point isn’t, ‘Is this AI as good as a psychiatrist with 30 years of experience?’” said Soutullo, who holds a John S. Dunn Professorship at McGovern Medical School. “The point is that both of them are doing different things, and they have strengths and weaknesses on both sides. Why pick one? You can use both.”
The team said they hope the system could eventually be used as an educational tool or as a supplement for early career clinicians or clinicians who don’t have psychiatry training.
“Just imagine you send this to a rural pediatrician who is seeing patients by themselves, and they can get this as an enhancement of their training,” Soutullo said. “That could be really good, because we could potentially detect symptoms that otherwise could be missed by clinicians who are not fully trained.”
The AI model was evaluated using video recordings of standardized patients undergoing mental status examinations, a mental health evaluation that serves a similar purpose as a physical exam. The standardized patients portrayed cases of schizophrenia, obsessive-compulsive disorder, and bipolar disorder at different severity levels.
The AI model, as well as teams of psychiatrists from UTHealth Houston and Yale, classified the standardized patients across 10 criteria: mood, appearance, behavior and cooperation, perceptions, speech, suicidality, the presence of delusions, obsessions or compulsions, and coherence and speed of thought processes.
The diagnoses from the psychiatrists were then compared with the AI model’s diagnoses.
The AI model was tested on its overall diagnosis as well as how it performed across each of the 10 individual criteria. While the tool made some mistakes on individual criteria, like appearance of the patient or the patient’s fine motor movements, the tool’s overall diagnosis was consistently accurate.
The team said the next step is to train the AI model to become more accurate at classifying individual criteria. For now, however, the tool has significant educational potential.
“It was equally good in observations, but not very good at some domains that involve the reasoning that a psychiatrist is doing,” said co-first author Benson Mwangi Irungu, PhD, assistant professor in the Department of Psychiatry and Behavioral Sciences at McGovern Medical School. “Even the AI’s mistakes could become teaching tools. Students could compare its assessments with those of experienced clinicians, examine where the reasoning went wrong, and learn how to avoid similar errors in their own clinical assessments.”
The research was funded in part by the New and Emerging Children’s Mental Health Researchers Initiative of the Texas Child Mental Health Care Consortium, and by the National Institutes of Health's Artificial Intelligence/Machine Learning Consortium to Advance Health Equity.
Additional authors with UTHealth Houston include Marsal Sanches, MD, PhD; Nurhak Dogan, MD; Pooja Chaudhary, MD, MPH; Mon-Ju Wu, PhD; Giovana Zunta-Soares, MD; and Jair C. Soares MD, PhD.
Andrés S Martin, MD, PhD, MPH, of Yale University is a co-author on the paper.