SYSTEM NOTICE

Auto translation by AI. Be sure, accuracy, nuances and authorial intent may not be fully reflected.
見出し画像

[Audio Commentary] Generalist Large Language Models (LLMs) Outperform Specialized Clinical AI Tools on Medical Benchmarks



In medical settings, specialized clinical AI assistants tailored for specific tasks are being rapidly adopted.
While these tools are often advertised as safer and more reliable than generalist LLMs, unlike state-of-the-art generalist models, they are rarely subjected to independent, quantitative evaluation, creating a significant evidence gap regarding their performance.
This situation can lead to avoidable clinical risks in settings where AI output influences diagnostic and management decisions.


To address this gap, a research team led by Krithik Vishwanath compared and evaluated two widely deployed clinical AI systems (OpenEvidence and UpToDate Expert AI) against three state-of-the-art generalist LLMs (GPT-5, Gemini 3 Pro, Claude Sonnet 4.5).
The evaluation used a 1,000-item mini-benchmark that combined MedQA, which measures medical knowledge, and HealthBench, which measures alignment with clinician judgment.


Surprisingly, the results of this evaluation revealed that generalist models consistently outperformed clinical tools. In particular, GPT-5 achieved the highest scores.


In terms of accuracy on MedQA, which tests medical knowledge, GPT-5 recorded 96.2%, outperforming clinical tools such as OpenEvidence (89.6%) and UpToDate (88.4%).
Furthermore, in the average consensus score of HealthBench, which evaluates alignment with experts, generalist models (average 91.7%) outperformed clinical tools (average 74.8%) by approximately 1.23 times, demonstrating a statistically significant advantage.

It was found that the clinical tools OpenEvidence and UpToDate Expert AI lagged behind in many clinically important aspects, such as completeness, quality of communication, context awareness, and reasoning for system-based safety.
In axis-level analysis, generalist models were superior in completeness, quality of communication, and context awareness.
Furthermore, in theme-level analysis, clinical tools scored lowest or near-lowest in categories such as urgent referrals and communication tailored to specialized knowledge.


These results suggest that tools currently on the market for clinical decision support may be lagging behind frontier LLMs. It has been pointed out that the Retrieval-Augmented Generation (RAG) mechanism heavily used by OpenEvidence and UpToDate Expert AI can impair model performance if incorrect information is retrieved or if it is not well-integrated with the foundation model.


As generative AI models are integrated into decision-making in medical settings, the discrepancy between advertised performance and actual performance poses avoidable clinical risks. Therefore, there is an urgent need for transparent, independent evaluations regarding their limitations, biases, and comparative performance before these technologies are introduced into patient-facing workflows.


References

Krithik Vishwanath, Mrigayu Ghosh, Anton Alyakin, M.S.E., Daniel Alexander Alber, B.S., Yindalon Aphinyanaphongs, M.D., Ph.D., Eric Karl Oermann, M.D. (2025). Generalist Large Language Models Outperform Clinical Tools on Medical Benchmarks. (Preprint)



*This article is AI co-created content. Content generation/summary: NotebookLM

いいなと思ったら応援しよう!

bycomet いつもありがとうございます。 これからも、どうぞよろしくお願いします。