Scientists have questioned the willingness of AI to help in healthcare
- Новости
- Science and technology
- Scientists have questioned the willingness of AI to help in healthcare
Over the past few years, artificial intelligence has made a real breakthrough in medicine. Modern neural networks have learned how to analyze X-rays, identify tumors, draw up medical reports, and even successfully pass exams that future doctors take. Against this background, predictions are increasingly being made that AI will be able to replace some medical specialists in the coming years. However, the scientists themselves urge not to rush to such conclusions. A new study has shown that high test scores of artificial intelligence (AI) do not mean that it can be trusted to treat real patients. About why medical neural networks are still only assistants to doctors and what obstacles stand in the way of their full implementation in healthcare, in the Izvestia article.
Why does an A on the exam still not make an AI a good doctor?
Over the past two years, large language models have repeatedly become the heroes of high-profile news. Some successfully passed the USMLE American medical exam, while others showed results comparable to the answers of experienced doctors to clinical questions. Such achievements have given rise to the feeling that artificial intelligence is almost ready to work with patients on its own. However, the authors of a new study in the journal Nature Medicine believe that this logic is becoming one of the main mistakes in evaluating medical AI today.
Researchers are paying attention to a problem they call "benchmark readiness" — in other words, successfully passing the tests does not mean that you are ready to work in a real clinic. Most modern assessments are based on pre-prepared sets of tasks where there is a correct answer. Real medicine works completely differently: patients rarely describe symptoms according to a textbook, diseases can mask each other, test results can be contradictory, and the doctor often has to make decisions in conditions of incomplete information.
The authors compare this situation with an exam for a pilot: a person can perfectly answer questions on aerodynamics and navigation, but this does not mean that he can handle landing an airplane during a thunderstorm. According to the researchers, something similar is happening with medical artificial intelligence: models perfectly solve standard tasks, but real clinical practice requires much more — the ability to work with uncertainty, unexpected circumstances and the unique characteristics of each patient.
That is why scientists propose to abandon the usual approach in which the success of AI is assessed only by the number of correct answers. Instead, they call for testing models in conditions as close as possible to the real work of a doctor: to evaluate how algorithms interact with medical staff, respond to incomplete data, explain their conclusions and behave in unusual situations.
The main problem is that medicine doesn't look like an exam.
One of the most serious problems of modern neural networks remains the so—called hallucinations - situations when the model confidently outputs information that does not correspond to reality. In everyday life, such a mistake can lead to an incorrect retelling of an article or a fictitious reference to a scientific paper. In medicine, the price of such inaccuracy turns out to be much higher: the algorithm is able to name a drug that does not exist, make a mistake in dosage, or refer to clinical recommendations that never existed.
Researchers call an equally serious problem the bias of the data on which the models are trained. If the algorithm mainly analyzed the medical records of patients of a certain age, gender, or origin, its accuracy may decrease markedly when working with other groups of people. That is why many experts believe that medical AI needs to be tested on as much diverse data as possible even before it is introduced into hospitals.
Another feature of artificial intelligence is that it rarely demonstrates insecurity. Unlike a doctor who can prescribe additional studies or refer a patient to a specialist, the language model often formulates the answer as if it is completely sure that it is right. According to the authors of the Nature Medicine reviews, it is this "false confidence" that is considered one of the most dangerous qualities of modern neural networks, since it can be difficult for a person to understand where the algorithm is really right and where it is only convincingly wrong.
Finally, medicine is not just about analyzing symptoms and making a diagnosis. The doctor has to take into account the patient's psychological state, explain difficult decisions to relatives, notice details that cannot be reduced to numbers, and take responsibility for the decision. So far, no language model is able to fully reproduce this aspect of clinical practice. That is why most researchers today consider artificial intelligence not as a substitute for a doctor, but as a tool that can help a specialist analyze information faster, but does not have to make final decisions on their own.
Why most studies don't prove anything yet
Despite the rapid development of medical artificial intelligence, scientists still cannot unequivocally answer the main question: do such systems really improve the health of patients? The reason is that most studies evaluate the performance of algorithms in artificial conditions — on exam questions, pre-prepared clinical cases, or sets of medical images. Only a small part of such developments reaches real practice.
This is confirmed by a large-scale review, also published in Nature Medicine earlier. The researchers analyzed 4,609 scientific papers on large language models in clinical medicine. It turned out that the vast majority of studies tested AI's abilities in simulations or theoretical scenarios. Real patients appeared in only about a quarter of the studies, and the authors found only 19 real randomized clinical trials — the gold standard of modern medicine.
According to the researchers, this situation is similar to testing a new drug exclusively in the laboratory: even if the drug works fine in a test tube, it is not enough to prescribe it to millions of people. He must go through several stages of clinical trials, prove safety and effectiveness, and only then get to hospitals. The authors believe that medical AI needs an equally rigorous verification system because the cost of making mistakes in healthcare is too high.
That is why in recent years, scientists have begun to develop special international recommendations for the introduction of artificial intelligence into clinical practice. One of these initiatives was the DECIDE-AI project, which offers uniform rules for evaluating medical algorithms before using them in hospitals. Researchers believe that neural networks should undergo step—by-step verification, from laboratory tests to work in real clinics under the constant supervision of specialists. In fact, we are talking about creating the same testing system for AI that already exists for new drugs and medical technologies.
Will AI be able to become a doctor one day
At the same time, scientists themselves do not consider artificial intelligence to be a dead-end technology at all. On the contrary, neural networks are already helping doctors find tumors on CT scans, analyze X-rays faster, identify signs of diabetic retinopathy, search for rare diseases, and automate the preparation of medical documentation. In many tasks, AI can significantly reduce information processing time and reduce the burden on specialists.
However, almost all modern research agrees on one thing: at the current stage of development, artificial intelligence should be considered primarily as a decision-making support tool, and not as an independent participant in the treatment process. The final diagnosis, the choice of treatment method and responsibility for the patient still remain with the doctor, who is able to take into account many factors inaccessible to the algorithm, from the features of the course of the disease to the emotional state of the person and the nuances of communication with his family.
The authors of a new article in Nature Medicine emphasize that the main task of the coming years is not to teach neural networks to answer exam questions even better, but to create new ways to assess their reliability in real medicine. Artificial intelligence must demonstrate not only high accuracy, but also resilience to incomplete data, the ability to interact correctly with doctors, explain their findings, and work safely in a wide variety of clinical situations.
The history of medicine has repeatedly shown that even the most promising technologies require lengthy testing before they become part of daily practice. Artificial intelligence is likely to follow this path too. Today, it is already able to significantly facilitate the work of doctors, but science, according to the researchers themselves, has not yet reached the point where the algorithm can be unconditionally trusted with human life. The high score on the exam turned out to be only the first step — the medical AI has a much more difficult test ahead: meeting with a real patient.
Переведено сервисом «Яндекс Переводчик»