Taboos of Generative AI: A Surgeon Explains 'What Information You Can and Cannot Input in Clinical Practice'
The other day, I saw a post on X where a doctor pasted test results into ChatGPT to consider a differential diagnosis. There was certainly no malice involved. However, do you know where the line is drawn between what is okay and what becomes a risk?
Just knowing that line will significantly change how you use AI from now on.
In this article, I will organize the basic boundaries you should keep in mind when using AI in clinical practice, as well as how to choose tools that you can start using today.
Generative AI does not 'know the answer'—can you explain the mechanism?
To describe how generative AI works in one phrase, it is 'a system that continues to select the next most probable word (token).' It learns from vast amounts of text data, such as internet articles, books, and papers, to generate a 'plausible continuation' that fits the context.
That is why it can converse naturally and answer medical questions fluently, but it does so not because it 'knows,' but because it is 'selecting plausible words.' This mechanism creates what is known as hallucination. It is more accurate to say it 'confidently makes mistakes' rather than 'tells lies.'
I am sure many of you have experienced this. When you ask ChatGPT to search for English papers, it returns papers that look 'plausible' in terms of title, author, and journal name, but they do not actually exist. This is hallucination.
Can you input patient information into AI?—Organizing laws, terms of service, and privacy
To conclude, as a general rule, you should avoid inputting patient personal information into standard cloud AI (ChatGPT, Claude) from the perspective of the Act on the Protection of Personal Information. However, there are gray zones, and the honest truth is that 'there are parts where legal interpretation has not yet been settled.'
Given this current situation, the starting point for the judgment I practice is to make it a habit to take one second to think, 'Can a person who reads this information identify the patient?'
Range of what you can input:
Checking general medical knowledge and guidelines
Educational questions that have been completely anonymized
Translation of literature and proofreading of English text (that does not contain patient information)
Organizing research ideas and considering document structure
Range of what you must not input:
Identifiers such as names, patient IDs, and dates of birth
Combinations of test values, vitals, and symptoms of a specific patient
Image data (CT, endoscopy, etc.)

The capabilities of three general-purpose AIs—ChatGPT, Gemini, and Claude: Which one should a doctor use?
I am increasingly being asked which AI should be used. I will summarize my impressions as of 2025–2026 (please check each company's website for the latest information, as performance and pricing are updated periodically).
I feel that the majority opinion is that Claude outputs the most natural Japanese among the three models. I also use it most frequently for summarizing Japanese text and writing documents.
Many people have the impression that Gemini is 'good at Japanese,' but this may be because Gemini's speed in processing long texts and its comprehensiveness make it appear 'good.' Gemini is useful for summarizing long documents.
If I had to choose just one, I feel that Claude is more user-friendly for doctors who do a lot of Japanese writing. However, honestly, there is no major difference, and the pros and cons change daily as models are updated. I believe the most important thing is to be able to use at least one of them correctly.

→Added on 2026/8/7
At this moment, I think the two leaders are ChatGPT and Claude! It is also attractive that these can easily use agent functions via ChatGPT Work and Claude Cowork.
Honestly, there is no major difference between these two (though each has its own strengths), so I think it is sufficient to use whichever one more people around you are using or whichever one you are interested in.
How to Use Medical-Specific AI: Options for Situations Where General-Purpose AI Cannot Be Used Directly
In situations where general-purpose AI cannot be used directly, there is the option of using AI designed specifically for medical purposes. Unlike general-purpose AI, a major feature is that it provides answers with literature citations.
MedGen
This is a medical-specific AI that is easy to use for checking Japanese guidelines and package inserts. It has strengths in searching for medical information in Japanese. There is no app; you log in and use it from a browser.
Referral code: S1XQ3D4I

OpenEvidence
This is an AI designed for medical professionals. It is supervised by the Mayo Clinic and also cooperates with the NEJM. All doctors can use it for free (uploading a medical license is required at registration; there are also cases where registration was possible with an employee ID).
It provides answers with citations based on papers and guidelines. Even if you ask in Japanese, it will answer in Japanese. However, since it focuses on English literature and overseas guidelines, there are limitations regarding Japanese insurance medical practice constraints and Japanese guidelines.
However, I want to emphasize that 'with citations' does not mean 'fact-checked.' You must verify the existence of the citations and judge the validity of the content yourself.
My own way of using it is mainly 'when I want to know about something outside my specialty.' In my own field of expertise, I can handle things by reading guidelines and key papers, but for fields completely outside my specialty, these tools are a convenient starting point. As support for daily clinical practice, I have the impression that the Japanese-specialized MedGen is easier to use as a reference.
(For those who want to dig deeper into literature searches, please also read the previous article, 'The Story of How a Surgeon's Literature Search Was Transformed by AI'.)https://note.com/yokubari_dr884/n/nd709ae05fefa)
AI That Runs Only on Your PC: Local LLMs and Hospital AI
The topic changes slightly here.
From what I have discussed so far, if you ask whether generative AI will never be introduced into medical records, that is not the case. The AI I have talked about so far involves uploading information to the cloud to process it.
In contrast, a local LLM (Large Language Model: an AI model specialized in understanding and generating text among generative AIs) is an AI that runs only on your own PC without connecting to the Internet. The information you input is not sent to external servers.
Recently, there has been an increasing number of cases where generative AI is being introduced into electronic medical records. There are mainly two types.
Local LLM type: Completed within the hospital server, patient information does not go outside
Encrypted cloud type: Data is encrypted and sent externally. A perspective on data processing outsourcing contracts is required
Just because you use generative AI in your electronic medical records does not mean it is safe to input patient information into generative AI on your personal computer; that is a major misconception.
An Action List You Can Start Today: Getting Started at Three Levels
You do not need to change everything difficult all at once.
Things you can do right now:
When asking generative AI anything related to clinical practice, make it a habit to pause for one second and ask yourself, 'Could someone reading this information identify the patient?' This will significantly reduce the risks of using generative AI.
Things you can do within this week:
Try using MedGen or OpenEvidence and input one question that has been on your mind from your recent clinical practice. Using AI to think about which parts of your daily work are 'verifiable tasks' is also effective for increasing the resolution of how you use these tools.
Things to work on next month:
If you want to utilize AI further, it might be a good idea to take a step beyond chat-based AI and try using autonomous AI agents like Claude Code or Codex.
(I plan to write more about this in the future.)
An Era of 'How to Use,' Not 'Whether to Use'
AI has already permeated both outpatient clinical settings and research environments. In reality, choosing 'not to use it' is becoming an opportunity loss. However, 'trying to use it for everything' will only pile up information security risks.
I feel that the middle ground—'understanding the mechanisms and using them'—which means being able to judge what information can be put into which AI, is the literacy required of doctors from now on.
If you were to change just one thing this week, what would it be? Even just 'checking for one second before inputting' is a sufficient first step. Let's start together.
Referral Code: S1XQ3D4I
Tags: #MedicalAI #Surgeon #ChatGPTUtilization #AIinMedicine #DoctorsUsingAI #OpenEvidence #MedGen #MedicalDX
