Which one to choose for academic use? Comparing 4 major AI services by price, features, and usability (March 2026 edition)
Comparison of ChatGPT / Gemini / Claude / Grok for researchers and clinicians
Practical perspective as of March 13, 2026
In recent years, LLMs such as ChatGPT, Gemini, Claude, and Grok have evolved rapidly, becoming deeply integrated into researchers' workflows, from paper writing, research proposals, and presentation materials to coding support.
On the other hand, what matters now is not just model performance.The differences in external information integration, file connectivity, and AI agent tools are now directly linked to the decision of "which one to use for daily tasks." era.
In this article, we compare the four services assuming academic use (paper writing, grant applications, conference presentation slide creation, and clinical question solving).
Conclusion first: Which one should you choose (in my case)
I will state my conclusion based on my current usage first.Please note that this includes a fair amount of personal (academia, rarely coding) preferences and financial sense (being frugal).
First, if you are only going to pay for one, ChatGPT Plus is a solid choice. It has the most balanced overall capability, covering search, research, document creation, and coding.
Beyond that, I use them differently as follows:
Japanese polishing, Claude in Chrome → Claude Pro
Images/Vision → Gemini (AI Plus)
Paper writing/Coding → Codex (included in ChatGPT subscription)
Grok → Wait and see
I will explain the detailed features and user experience of each service in the main text.
Please refer to the figure below for the main points.


1. Comparison of prices and plans (for individuals)
What changed significantly as of March 2026 is that each company has increasingly polarized into low-cost plans and premium plans. Google has newly established AI Plus, and OpenAI has added Go.
ChatGPT (OpenAI)
Free Plan: Primarily limited access to GPT-5.3.
Go (approx. $8/month): Low-cost plan. Access volume to GPT-5.3 and message limits are expanded compared to the free plan. However, I think it's insufficient for professional work.
Plus ($20/month): The go-to for individuals. It includes GPT-5.4 Thinking, deep research, agent mode, Projects, and Codex, offering a good balance.
Pro ($200/month): GPT-5.4 Pro, unlimited access, priority speed Codex, etc. For heavy-duty work.
Gemini (Google)
Free Plan: In addition to 3 Flash, you can use variable access to 3.1 Pro, Deep Research, Gemini Live, Canvas, and Gems. The free tier is quite generous.
Google AI Plus (¥1,200/month): A newly established low-cost plan. Includes enhanced access to 3.1 Pro, Nano Banana Pro-series image generation, and Workspace benefits.
Google AI Pro (¥2,900/month): 3.1 Pro, 1 million token context, 2TB storage, and higher-tier Code Assist / CLI access.
Google AI Ultra (¥36,400/month): The top tier. Includes Deep Think, Gemini Agent (English only), etc.
Claude (Anthropic)
Free Plan: Limited. You can create up to 5 Projects even on the free plan.
Pro ($20/month / $200 annually): The standard paid plan.
Max 5x ($100/month): For those who want more headroom than Pro. It might be tough to use as your main tool, including Claude Code, without subscribing to at least this.
Max 20x ($200/month): For those using it for full-scale operations, including Claude Code.
Grok (xAI / X)
Free: Grok 4.1 is currently rolling out to all users.
SuperGrok ($30/month): Standalone subscription on grok.com.
SuperGrok Heavy ($300/month): Access to Grok 4 Heavy.
X Premium+ (¥6,080/month): A path to use Grok via X.
2. Features and user experience of each model
ChatGPT
The current mainstays are GPT-5.4 Thinking and the higher-tier GPT-5.4 Pro. I don't trust GPT-5.3 Instant very much.
GPT-5.4 Thinking has had its context length expanded to a total of 256k tokens, which has significantly improved long-form reasoning and the retention of long-text context. For drafting papers, grant applications, research, and creating structured documents, its stability in "producing well-thought-out text" is a cut above other models.
Incorporating external information is also a major strength of ChatGPT. Since it composes text while naturally weaving in search results, it is highly compatible with the typical research workflow of verifying evidence -> refining -> formatting. When combined with deep research or agent mode, you can run the entire process from information gathering to drafting in one go.
On the other hand, Japanese can be a bit quirky, to be honest. I think English documents are at a level where they can be used as-is (if you give good instructions), but in situations where "easy-to-read Japanese" is required, such as in Japanese review articles or Note articles, I often pass them to Claude Opus for editing.
By the way, I am currently subscribed to Pro ($200/month), but I don't use the web version of Pro that often, so I might switch back to Plus. However, currently, the Codex limit is doubled, so if I start hitting the Codex limits more often, I might consider the Pro subscription.
Gemini
I feel improvements with Gemini 3.1 Pro, but looking at the model alone, it is undeniable that it falls short in practical use compared to GPT-5.4 Thinking and Claude Opus 4.6. Even for long-form tasks, which used to be Gemini's strong suit, I get the impression that its advantage is disappearing due to the performance improvements of other models.
It is not good at search-related tasks, and I have the impression that there are many hallucinations. There are also times when it refuses to answer unnaturally, and there are many situations where the behavior of the web version feels unstable. The fact that the workflow of working while searching doesn't feel as natural to me as it does with ChatGPT is the biggest reason why my usage frequency has dropped.
However, vision (image recognition) is still good. For image generation, I think Nano Banana is easy to use with few garbled Japanese characters. It still has its uses for figures and images in slides.
Also, I often use it for drafting Gmail. The fact that it is quite usable even for free is a genuinely good point.
Furthermore, I switched my Gemini subscription from Pro to AI Plus (¥1,200/month). I wasn't using it as frequently as Pro, so this is sufficient in terms of cost.
Claude
Opus 4.6 is the star. While I get the impression that its overall intelligence is slightly behind ChatGPT, Claude still has the edge in the naturalness of Japanese and the polish of its writing style. In situations where "easy-to-read Japanese" is required, such as Japanese review articles, conference abstracts, and Note articles, the method of polishing drafts created with ChatGPT using Claude remains effective.
In the paid plan, you can use Research, and it can also link with Google Workspace's Gmail / Calendar / Drive. Its "ability to fetch external information" has definitely improved compared to before. However, my impression that ChatGPT is a cut above when it comes to finalizing the output remains unchanged.
Claude in Chrome is personally very useful to me. It handles routine input tasks on browsers, such as registration work on journal sites when submitting papers. I intend to keep paying for Claude Pro ($20/month) for a while until ChatGPT's Agent can provide an equivalent UX.
以前からずっとAIにやってほしかった論文の共著者の登録ですが、Claude in ChromeでOpus4.6を使ってようやく完璧にできました!
— 限界助教|ChatGPT/Claude/Geminiで論文作成と科研費申請 (@genkAIjokyo) February 6, 2026
多施設研究の20+の共著者の名前、所属の入力を完璧にこなしてくれました
データはある程度成型して渡してあげるのがコツっぽいです https://t.co/af76Wp88r0
When doing substantial work on the Pro plan, I get the impression that it's relatively easy to hit the limit on both the Web version and Claude Code. Usage is counted collectively across Web, Desktop, Mobile, and Claude Code, so if you are using it heavily, the Max subscription becomes a consideration.
Grok
For consumers, Grok 4.1 is widely available, and there is Grok 4 / Grok 4 Heavy for higher-tier subscribers. While it feels less restrictive and its integration with X and real-time web information is distinctive, it still needs more maturity in terms of structural capability for research documents and clinical use, as well as stability in Japanese. My current impression is that it is not yet a viable replacement for ChatGPT, Claude, or Gemini for academic purposes.
3. Comparison of external files, development environments, and integration features
The difference is now significant not only in model intelligence but also in **"whether it can output files usable in practice," "whether it connects to development environments," and "whether it can incorporate external information with citations."**
3-1. Input/Output of external files (Office / PDF, etc.)
ChatGPT is quite well-equipped for documents. You can export from Canvas in PDF, DOCX, or Markdown, and Data Analysis can read Excel, CSV, PDF, JSON, etc. It can also handle Google Drive / OneDrive files. However, it is not as direct an entry point for "delivering" Office files as Claude; it involves combining Canvas, Data Analysis, agents, and Codex.
Gemini is designed to create within Google Workspace rather than directly exporting Office-compatible files. You can export from the Gemini app to Google Docs or Gmail, and generate or edit slides within Slides. It is more natural to create in Docs, Slides, or Sheets and convert if necessary, rather than outputting Word or PowerPoint directly.
Claude is the most straightforward in this area. It can directly create and edit xlsx, pptx, docx, and PDF files, which can then be downloaded or saved directly to Google Drive. Recently, enhancements to Claude in Excel and a research preview of Claude in PowerPoint have also emerged. As of March 2026, Claude remains the best for ease of Office delivery.
ClaudeのWord編集Skill使った校正が実用的!
— 限界助教|ChatGPT/Claude/Geminiで論文作成と科研費申請 (@genkAIjokyo) February 21, 2026
変更履歴つけて理由をコメントでつけられます
→承認で変更完了!
(ChatGPTは変更履歴はできたりできなかったり)
prompt:
添付のワードの英文校正をしてください。修正履歴がわかるように修正をして修正理由を日本語コメントで指摘してください。 pic.twitter.com/XPdhpFDDYg
Grok currently has a weaker appeal for directly delivering Office artifacts. Grok Business / Enterprise has Google Drive integration, allowing it to answer while referencing files on Drive, but it lags behind the other three in terms of document delivery capabilities for individuals.
3-2. Context management: Projects / Gems / Knowledge
ChatGPT Projects has become quite polished as a workspace for long-term projects. It can now incorporate not just files and instructions, but also Slack and Google Drive links, past chats, and even snippets of notes as project knowledge.
Claude Projects is now available to free users for up to 5 projects, and the paid version is designed to expand project knowledge via RAG up to 10 times the scale. It remains well-suited for long-term work that requires accumulated knowledge.
Gemini does not have a front-facing concept like Projects, but it can be largely substituted with Gems. These are custom Gemini instances with fixed roles, rules, and prerequisite materials, making them convenient for creating templates for repetitive tasks.
About NotebookLM
As a tool related to Gemini, I should also mention NotebookLM. It is a tool that allows you to upload materials and ask questions or get summaries based on their content, and it is often a topic of discussion.
However, to be honest, I hardly use it myself. There are three reasons for this.
First, if I'm going to launch NotebookLM, I often find it easier to just throw the materials directly into Gemini. The step of creating a notebook feels like an extra hassle in my workflow.
Second, I don't really use it in a way where I repeatedly ask questions or extract information from the same material. I believe that is where the true value of NotebookLM lies, but there are few such situations in my work.
Third, when I want to write something based on context, I feel that keeping notes in Obsidian and combining them with Codex has more potential for development. This is purely a personal preference, but I believe that a style of accumulating knowledge in a local Vault is more reusable in the long run.
I don't think NotebookLM itself is a bad tool, but for the reasons mentioned above, it is currently not integrated into my workflow.
3-3. Coding Assistance / Agentic IDE
My current workflow: Codex (Desktop App)
I have currently settled on the OpenAI Codex desktop app. It was initially only for Mac, but it recently became available for Windows as well.
The good thing about Codex is, first of all, that you almost never hit the limit even on the Plus plan. With Claude Code, there are times when I worry about the shared usage cap, but with Codex, that is rarely a concern. For those who write code daily, being able to use it freely without worrying about limits is a major benefit.
Another point is that Codex's web search is as high-performing as ChatGPT's, allowing you to complete error resolution and specification checks within the tool. Since Codex is included in Plus / Pro, no additional payment is required.
On the other hand, a characteristic of Codex is that it only says what is absolutely necessary. There is no small talk or explanation; it returns only the requirements. I have heard people say it is "hard to talk to," but since it is a work partner for me, I don't mind that.
It is also possible to write academic papers with Codex.
About Claude Code
Claude Code is also a good tool. As mentioned earlier, since usage is counted across the Web version, Desktop, Mobile, and Claude Code, you would want the Max plan if you intend to run it autonomously for long periods.
My impression of Claude Code is that it feels like the model's performance is well-complemented by the harness (the tool's internal mechanism). In addition, it has a UI and operational feel that hits the spot, so it is understandable why it has many fans. It shows the typical politeness of Claude in design and text-related tasks, so I think it is suitable for those who want to use it for a wide range of tasks beyond just coding.
As for me, even when I write code, it is only for statistical analysis, and I do not build complex programs. Therefore, I honestly cannot judge which is better between Codex and Claude Code for my purposes. However, looking at the model alone, I have the impression that Codex is more likely to handle difficult tasks.
About Google Antigravity (Agentic IDE)
Google is also rolling out Antigravity alongside Gemini Code Assist / Gemini CLI. The AI Pro / Ultra plans include increased rate limits as a benefit.
Antigravity was quite usable even for free at the beginning. When Claude Opus 4.6 was relatively available in large quantities, I used it quite a bit as well. However, perhaps because many users were abusing it, the rate limits seem to be quite strict now. Since I also downgraded my Gemini plan, I haven't been using it much lately either.
4. About Presentation Creation
Each company is making progress in slide creation. ChatGPT is promoting slideshow creation in agent mode, Claude supports direct generation of pptx files, and Gemini can also generate and edit slides within Google Slides.
As a specialized slide tool, Genspark is well-made; it can pull information from the web to create an outline and export it to .pptx / PDF / Google Slides. The quality of the finish is certainly high. However, since Genspark has almost no use cases other than slides, it is questionable whether it is worth paying for just that.
My slide creation method: Codex or Claude Code's "Skills"
I personally use Codex or Claude Code and utilize the Skills mechanism to generate .pptx files.
Claude Code has better design skills. It has a better sense of color and layout, and it produces good-looking slides even with minimal instructions. On the other hand, with Codex, you can also create sufficiently good slides by restricting templates and design rules in your instructions.
For more details on this method, please refer to the article below.
I believe it is more rational to use the pptx generation features of ChatGPT (Codex) or Claude, which you are already paying for, rather than paying separately for Genspark just for slides.
5. Summary: My current workflow
Drafting, structuring, and research: ChatGPT
With its comprehensive capabilities including search, deep research, Projects, Canvas, and Codex, ChatGPT is the easiest to use for creating outlines. If you are only going to pay for one service to start, I still think ChatGPT Plus is the solid choice.
Japanese polishing and style adjustment: Claude
Claude is still superior in terms of natural Japanese, the warmth of the writing, and the sophistication of the expressions. The workflow of polishing drafts created in ChatGPT using Claude remains effective. I also find Claude in Chrome useful for tasks like registering for paper submissions, so I intend to continue my Pro subscription for a while.
Vision and image generation: Gemini
Gemini still has its place for diagrams, image recognition, and visual processing. Nano Banana-based image generation is quite useful when the use case fits. However, I have switched from Pro to AI Plus.
Coding and analysis: Codex
When it comes to research code, verification, and library research, Codex suits me best. I also use Claude Code as a second opinion, but my basic approach is to start with Codex.
Slide creation: Codex or Claude Code (utilizing Skills)
I generate pptx files using Skills in Codex / Claude Code. Claude Code has better design, while Codex is perfectly capable if you provide template instructions.
Final organization
As of March 13, 2026, my view is that if you have to choose just one for academic or professional use, it is ChatGPT Plus. Beyond that,
Japanese polishing, Claude in Chrome → Claude Pro
Images/Vision → Gemini (AI Plus)
Coding → Codex (included in ChatGPT subscription)
Grok: wait and see
I use them in the following ways.
Rather than "trying to do everything with one tool," I feel that using ChatGPT as the core and adding other tools only for necessary areas is the most reasonable way to use them for now.
