What Makes an AI 'Easy to Use' for Doctors? Thoughts Based on a Paper Comparing Medical-Specialized AI vs. General-Purpose AI
I recently read a paper comparing "medical-specialized AI," such as OpenEvidence and UpToDate Expert AI, with general-purpose LLMs like the GPT series. When evaluating the quality of responses to USMLE-style questions and clinical cases, the results showed that, surprisingly, general-purpose LLMs scored better than medical-specialized AI.
The paper is here: https://arxiv.org/pdf/2512.01191
Note that this paper is a preprint posted on arXiv and has not yet undergone peer review. Keeping that in mind, I think it is important to consider how to interpret what is written.
Medical-Specialized AI vs. General-Purpose LLMs: A Brief Overview of the Paper
The paper roughly compares the following:
Medical-Specialized AI
・OpenEvidence
・UpToDate Expert AI
General-Purpose LLMs
・GPT-5
・Gemini3pro
・Claude, etc.
The evaluation used benchmarks that looked at the quality of responses to clinical cases from multiple perspectives, not just USMLE-style exam questions. The metrics examined include items such as the following:
Accuracy
Completeness
Contextual understanding
Quality of communication
Adherence to instructions, etc.
The results showed that general-purpose LLMs outperformed medical-specialized AI in many areas. Of course, this is only within the context of benchmarks and cannot be directly applied to clinical practice as is. Even so, it is a result that makes one want to rethink the intuition that "AI built specifically for medicine must be stronger in a clinical setting."
This paper led me to think, "So, what exactly is an 'easy-to-use AI' for a doctor?" From my experience using AI tools between clinical duties and other work, I feel that it is not just about "intelligence," but that certain "personality" traits are quite important.
Conditions for ease of use that cannot be measured by "intelligence" alone
First, there is the question of "whether it can fail safely." The scariest thing for a doctor is an AI that is confidently wrong.When plausible-sounding terminology and paper titles are presented, it becomes harder for humans to apply the brakes. On the other hand, an AI that properly draws lines at its limits—such as by saying, "Please check the guidelines directly from here on," or "Evidence is divided in this area, so there is no single conclusion"—is a partner I can still work with, even if it makes some mistakes.
Another point is whether it aligns with how a doctor organizes their thoughts.An AI that can reorder a differential diagnosis into categories like "common conditions," "conditions you don't want to miss," and "rare but lethal conditions" is far easier to use than one that just tries to guess the diagnosis in one shot. Regarding treatment plans, if it can help organize the standard guidelines alongside the reasons for deviating from them or the room for flexibility based on patient values, it begins to function as a "consultation partner."
Furthermore, how much it helps with explaining things to patients is also important. An AI that can adjust the tone and length of an explanation based on requests like "explain it so a grade-schooler can understand," "for someone who has just been told they have cancer," or "in a format easy to share with elderly family members" becomes a very reassuring presence in the outpatient clinic. While you ultimately need to put things into your own words, having a draft versus having to come up with an explanation from scratch makes a huge difference in terms of time and mental energy.
And finally, although it is a practical matter, there is the issue of how naturally it can blend into the workflow.No matter how smart it is, if you have to open a separate PC, log in, and copy-paste from the medical record every time, it will never take root in a busy clinical setting. Mundane usability—such as being able to call it up naturally while writing medical records or being able to immediately pour content into a draft for a referral letter or summary—is what ultimately determines the success or failure of a service.
Ten years ago, when I started Antaa and was building products, the point I was particularly conscious of was exactly this: 'Is it designed into the daily workflow?' Whether Antaa's product could be gently placed as an extension of the very natural behavior of doctors searching for cases and information on their smartphones daily. It was about not forcing new behavioral patterns, but placing it within existing movements.
I think this feeling also applies when thinking about AI tools.
What kind of AI will 'survive' in medicine from now on?
When I organize it this way, I feel that the AI that will survive in medicine from now on will be selected not by the label of 'medical-specific or general-purpose,' but rather by whether it can satisfy the following conditions.
Being able to properly say 'I don't know' when it doesn't know.
Being able to organize differential diagnoses and policies in a way that matches the doctor's thought process.
Being able to lighten the burden of verbalization and documentation for doctors, such as medical record summaries or drafts of referral letters.
Being able to naturally integrate into the flow of daily work without forcing new lines of movement.
Furthermore, it will also become important in the long run that the line between what the AI handles and where the human takes responsibility is designed into the product from the start.
Rather than 'choosing' AI tools, I think only those that possess these conditions will ultimately remain in the medical field. In that process, I feel that doctors will also be required to articulate 'how much can be entrusted' and 'what parts I will continue to handle myself,' and adjust these with the people building the products.
There is also an audio episode where I talked about this theme. If you would like to listen to it, please click here.
