Scientists compared the work of AI and psychiators in assessing mental state
- Новости
- Science and technology
- Scientists compared the work of AI and psychiators in assessing mental state
Researchers from the University of Texas at Houston and Yale University (the organization's activities are considered undesirable in the Russian Federation) compared a multimodal model of artificial intelligence with the assessments of psychiatrists during a mental status examination. The system analyzed video recordings of standardized patients and evaluated 10 parameters of mental state. The "Izvestia" article describes how the AI coped with this assessment and what limitations the study revealed.
How did the AI assess the mental state
The work was carried out by specialists from the University of Texas at Houston together with researchers from Yale University. The results are published in the scientific journal npj Mental Health Research (MHR). The scientists tested the multimodal Qwen3-Omni model, which is capable of simultaneously processing video, images, sound and text.
The study used video recordings of standardized actor patients who reproduced clinical cases of schizophrenia, obsessive—compulsive disorder, and bipolar disorder. Different levels of symptom severity were presented for each diagnosis. In total, the researchers compared 396 classifications according to 10 parameters of the mental status examination.
AI and two groups of psychiatrists from the University of Texas at Houston and Yale University evaluated the same materials. Separately, the researchers compared the results of the experts themselves to determine the degree of consistency of their assessments.
What the model analyzed
The examination assessed the patient's mood, appearance, behavior, and cooperation with the doctor, as well as perception, speech, thoughts of death, delusions, obsessions, and compulsions, as well as the coherence and speed of thought processes.
The model could take into account information from multiple visits to a single patient. Based on this information, the system generated assessments of individual mental state parameters and written explanations of the results.
The results were then compared with the psychiatrists' estimates. This allowed us to determine which characteristics the model recognized with greater or lesser accuracy and in which cases its conclusions differed from the expert ones.
To what extent did the results of AI coincide with the assessments of psychiatrists
The consistency between the two groups of psychiatrists was high: Gwet's AC1 score was 0.87. When comparing the model with experts, it ranged from 0.70 to 0.72, which corresponds to moderate consistency.
The results varied depending on the parameter being evaluated. The model performed better with characteristics that could be directly observed from video or determined from speech. There were more difficulties in assessing conditions that required interpretation of the patient's subjective experience, such as delusions or perceptual impairments.
With more severe symptoms, the consistency of the model with expert estimates also decreased. Thus, the ability of the system to assess the mental state depended both on a specific parameter and on the severity of symptoms.
In which cases did the model make mistakes more often?
The researchers separately examined situations where the psychiatrists' assessments coincided or differed. According to the agreed estimates of specialists, the model was wrong in about 17% of classifications. When the experts differed among themselves, the error rate of the system increased to 41-59%, depending on the group of specialists.
According to the authors, this shows that some of the discrepancies are related not only to the work of the model itself, but also to the difficulty of interpreting psychiatric symptoms. At the same time, the AI indicators remained below the degree of consistency between the two groups of psychiatrists.
What limits the use of AI
The study used standardized patients, rather than people undergoing psychiatric evaluation in routine clinical practice. This approach has made it possible to control diagnoses and severity of symptoms, but limits the portability of the results to real medical conditions.
In addition, the researchers studied one model, Qwen3-Omni, and data obtained from two medical centers. Therefore, the results cannot be automatically distributed to other artificial intelligence systems or all types of psychiatric examinations.
The authors consider the development as a tool that can be used in conjunction with specialists, and not as an independent replacement for a psychiatrist. Additional studies will be required to evaluate its performance in a clinical setting.
Переведено сервисом «Яндекс Переводчик»