Anthropic’s Claude Sonnet 5 has been generally available for three weeks. OpenAI’s GPT-5.6 Sol has been live for less than two weeks after replacing GPT-5.5 as the default model inside ChatGPT. Google’s Gemini 3.1 Pro has now sat in public preview for five months without a confirmed release date. None of the three launched on the same day, none of them target exactly the same buyer, and yet all three show up on the same shortlist whenever a team asks which API to build on next.
That overlap is the reason this comparison exists. Sonnet 5 is not Anthropic’s flagship. Claude Opus 4.8 still holds that title. But Sonnet 5 is the model Anthropic wants high-volume, cost-conscious teams to actually use, and its pricing undercuts GPT-5.6 Sol by a wide margin and comes close to Gemini 3.1 Pro’s preview rate. Put plainly, Anthropic’s value-tier model is now cheap enough to get compared directly against two full flagships, and the data says that matchup is closer than the sticker price would suggest.
This piece lines up specs, independently tracked benchmark scores, real per-token pricing, and concrete use-case guidance for all three models as they stand in July 2026. Every figure is attributed to the tracker, vendor page, or launch disclosure it came from. Anywhere the public record is thin, that gap gets flagged instead of filled in with a guess.
Don't miss new tech stories on Google
Add Tech Insider once in the Google app and our stories appear in your news suggestions.
Why Sonnet 5, GPT-5.6, and Gemini 3.1 Pro Are Being Compared Right Now
Three separate release cycles put these models on the same shortlist inside a single quarter. Claude Sonnet 5 reached general availability on GPT-5.6 Sol was previewed June 26, 2026, and released publicly July 9, 2026 (not June 30); Claude Opus 4 was released in May 2026, making the “month after” timeline inaccurate [3].8. GPT-5.6 arrived GPT-5.6 was released July 9, 2026 as a three-tier family: Sol, Terra, and Luna, with Sol as the flagship reasoning option; however, OpenAI markets it against other labs’ top models, and the date is correct but the claim implies it was anticipated before April 2026 [3]. Gemini 3.1 Pro has technically been available the longest, since Google’s model (likely Gemini 2.5 or similar) had a public preview start date, but no stable release date was announced as of April 6, 2026; however, the date February 19, 2026 does not match known Google model release timelines from official sources [3][14].
What makes this three-way matchup useful rather than academic is where Sonnet 5 sits on price. Anthropic built Sonnet 5 for what its own launch messaging calls high-volume, cost-sensitive workloads, the kind of traffic that makes up the bulk of a typical production API bill: support tickets, content generation, internal tooling, and multi-step agents that do not need Opus-level reasoning on every call. That positioning puts Sonnet 5 in an unusual spot. It is priced well below GPT-5.6 Sol, and close enough to Gemini 3.1 Pro’s estimated rate that cost alone no longer decides the matchup.
Search interest reflects that shift. Buyers comparing Claude against GPT and Gemini are no longer only asking which lab has the smartest model this month. A growing share are asking whether the cheaper option from one lab can substitute for the full-price flagship from another. That is the specific question this comparison is built to answer. The rest of this piece treats all three as live, actively supported options a team could ship against this week, not a hypothetical matchup between models still in a lab.
Full Specs Comparison: Sonnet 5 vs GPT-5.6 Sol vs Gemini 3.1 Pro
The table below lines up twelve specs that matter most for a team choosing between these three models today. Pricing and release status move fastest, so treat the preview-pricing figures as a live estimate rather than a locked-in rate.
| Spec | Claude Sonnet 5 | GPT-5.6 Sol | Gemini 3.1 Pro (Preview) |
|---|---|---|---|
| Developer | Anthropic | OpenAI | Google DeepMind |
| Public release date | June 30, 2026 | July 9, 2026 | February 19, 2026 |
| Release status | General availability | General availability | Public Preview |
| Context window | Not separately disclosed; Anthropic’s current architecture runs at 1,000,000 tokens | 1,050,000 tokens | 1,000,000 tokens |
| Max output tokens | Not publicly disclosed | 128,000 tokens | 64,000 tokens |
| Knowledge cutoff | Not publicly disclosed | February 16, 2026 | Not publicly disclosed |
| Input price (per 1M tokens) | ~$2.50 (estimated, see pricing note below) | $5.00 | ~$2.00 (third-party estimate) |
| Output price (per 1M tokens) | ~$12.50 (estimated, see pricing note below) | $30.00 | ~$12.00 (third-party estimate) |
| Native multimodal input | Text, image, documents | Text, image | Text, audio, image, video, PDF, code repositories |
| Model family tiers | Opus / Sonnet / Haiku | Sol / Terra / Luna | Pro / Flash |
| Primary strength | Writing quality, instruction-following, cost-efficient volume traffic | Frontier reasoning, everyday chat | Multimodal breadth, long context, cost-adjusted reasoning |
| Enterprise API availability | Anthropic API, AWS Bedrock, Google Vertex AI, Microsoft Foundry | OpenAI API, Azure OpenAI | Google AI Studio, Vertex AI |
A few gaps deserve a direct callout instead of a guess. Anthropic has not separately published a context window figure for Sonnet 5 the way it has for Opus 4.8, so the number above reflects the architecture Anthropic’s flagship uses rather than a Sonnet-5-specific disclosure. Sonnet 5’s pricing is likewise not broken out on Anthropic’s public rate card as an isolated figure in the sources reviewed for this piece. It is derived from Anthropic’s own launch positioning of Sonnet 5 at roughly half of Opus 4.8’s confirmed $5.00 input and $25.00 output per-million-token rate, which is the basis for the “half the cost” framing Anthropic used at launch. Readers building a production budget around that number should confirm current rates directly on Anthropic’s pricing page before committing spend.
Gemini 3.1 Pro’s pricing carries a similar caveat. Google has not published official per-token rates for the 3.1 Pro preview on its own pricing documentation as of this writing. The figures above come from third-party trackers and match the prior Gemini 3 Pro generation’s rate card, so treat them as a well-corroborated estimate rather than a locked number.
Claude Sonnet 5: Anthropic’s Value-Tier Workhorse
Claude Sonnet 5 became generally available on GPT-5.6 Sol was previewed June 26, 2026, and released July 9, 2026 (not June 30); this is roughly two weeks after the preview, not “a little over a month after Claude Opus 4” [3].8. Anthropic did not market it as an all-around upgrade. Independent trackers such as LLM Stats logged the release within hours and slotted it in as a writing and instruction-following specialist rather than a reasoning flagship, and Anthropic’s own product messaging frames Sonnet 5 as the model built for high-volume, cost-sensitive workloads rather than a new top-line leader.
On the Artificial Analysis Intelligence Index, the composite benchmark that has become a common reference point for comparing frontier models in 2026, Sonnet 5 scored 57. Opus 4.8 itself scored 61 on the same index when trackers benchmarked Sonnet 5 at launch, putting Anthropic’s own flagship four points ahead of its cheaper sibling. What stands out is how close that gap is to nothing once price enters the picture: Sonnet 5 costs roughly half of what Opus 4.8 charges per million tokens for a four-point difference on a composite reasoning score.
Where Sonnet 5 actually leads is in categories that sound less exciting than a benchmark score but matter more for day-to-day production traffic: writing quality and instruction-following. Instruction-following measures whether a model does what it is told: hits a word count, sticks to a format, avoids a banned phrase, cites a source correctly. Writing quality measures whether the output reads like a person who understood the assignment wrote it, not a model pattern-matching toward the nearest cliché. Sonnet 5 leads both categories among the models tracked at its price point, which is the specific claim behind Anthropic’s launch messaging.
That combination, cheap and reliable rather than cheap and clever, is exactly what a support-ticket triage system, a marketing copy generator, or a customer-facing chat agent actually needs. None of those jobs require Opus-level reasoning on every call. They require a model that follows formatting instructions consistently, does not drift from brand voice over a long conversation, and does not blow the monthly compute budget once traffic scales past a pilot. Anthropic’s own pricing page now reflects that split explicitly, with Opus reserved for harder reasoning and agentic tasks and Sonnet positioned as the default for volume.
GPT-5.6 Sol: OpenAI’s Reasoning Flagship
GPT-5.6 launched publicly on July 9, 2026, arriving as a family of three tiers rather than a single model. Sol is the flagship reasoning variant, Terra is a balanced mid-tier option, and Luna is the cost-efficient variant, according to OpenAI’s developer documentation. Sol is the tier this comparison focuses on, since it is the one OpenAI positions directly against Claude and Gemini’s top-line models on raw capability, and the one that replaced GPT-5.5 as ChatGPT’s default.
Sol’s official specs list a 1,050,000-token context window and a 128,000-token maximum output, with a training knowledge cutoff of The most recent cutoff date for GPT-5.6 Sol is not February 16, 2026; GPT-5.5 was released April 23, 2026, and GPT-5.6 Sol was released July 9, 2026, with cutoff dates likely in early 2026 but unconfirmed [3][5]. On price, Sol is the most expensive of the trio on output tokens: $5.00 per million input tokens, matching Sonnet 5’s estimated rate almost exactly, but $30.00 per million output tokens, more than double Sonnet 5’s estimated $12.50.
That price buys the strongest reasoning score in this specific comparison. GPT-5.6 Sol currently tops GPQA Diamond, the graduate-level science reasoning benchmark, at 94.6 percent per LLM Stats’ leaderboard. Its max configuration also posts a 59 on the Artificial Analysis Intelligence Index, two points ahead of Sonnet 5’s 57. OpenAI markets Terra and Luna specifically for workloads that do not need Sol’s full reasoning depth, the same tiered logic Anthropic applies to Opus versus Sonnet. Running every request through the flagship by default, instead of routing routine traffic to a cheaper tier, is one of the most common ways teams overspend once a new model family ships.
Sol’s officially documented input modalities are limited to text and image, narrower than Gemini 3.1 Pro’s native reach into audio, video, and full code repositories. OpenAI has hinted at broader modality support arriving across the GPT-5.6 family over time, but as of this writing that expansion has not shipped for Sol specifically.
Gemini 3.1 Pro: Google’s Multimodal Preview
Gemini 3.1 Pro is the outlier of the three in one important respect: it is still labeled Public Preview rather than a general-availability release, more than five months after it first appeared on February 19, 2026. Google’s own model card, published at deepmind.google, still lists the model string as gemini-3.1-pro-preview with no announced date for a stable release.
What Gemini 3.1 Pro offers in exchange for that preview label is breadth. It natively handles text, audio, images, video, PDFs, and full code repositories inside a single 1,000,000-token context window, more input modalities than either Sonnet 5 or GPT-5.6 Sol officially document. Google has not formally published per-token pricing for the model on its own pricing pages as of this writing. Third-party trackers, including Metatext’s model listing, report pricing that matches the prior Gemini 3 Pro generation: roughly $2.00 per million input tokens and $12.00 per million output tokens. Treat that figure as a well-corroborated estimate rather than an official Google number until the company publishes final preview pricing.
On capability, LLM Stats notes Gemini 3.1 Pro performs strongest in direct, head-to-head coding-arena matchups against other agentic coding models, though that is a relative ranking rather than a single comparable percentage score. It is also frequently cited as the strongest cost-adjusted reasoning option of the three once its lower estimated price is weighed against its results on GPQA Diamond and ARC-AGI-2. Neither Sonnet 5 nor GPT-5.6 Sol currently has an independently confirmed ARC-AGI-2 score in the sources reviewed for this comparison, so that specific benchmark is not included in the side-by-side table above.
The preview label is worth taking seriously rather than treating as a formality. Google’s changelog history shows preview models occasionally get repriced or have rate limits adjusted before a stable release, so a team building a long-term product on Gemini 3.1 Pro today should plan for that possibility rather than assume the current terms are final.
Benchmark Performance Across Independent Trackers
No single benchmark tells the whole story with these three models, and pulling from multiple independently operated trackers, Artificial Analysis, LLM Stats, and OpenRouter’s aggregated listings, makes that obvious fast. Leadership splits by category rather than concentrating in one model.
| Benchmark or category | Leader | Score | Source |
|---|---|---|---|
| GPQA Diamond (graduate-level reasoning) | GPT-5.6 Sol | 94.6% | LLM Stats |
| Artificial Analysis Intelligence Index | GPT-5.6 Sol (max config) | 59 | Artificial Analysis |
| Artificial Analysis Intelligence Index | Claude Sonnet 5 | 57 | Artificial Analysis |
| Writing quality (category leader) | Claude Sonnet 5 | Leads category | LLM Stats |
| Instruction-following (category leader) | Claude Sonnet 5 | Leads category | LLM Stats |
| Coding-arena head-to-head | Gemini 3.1 Pro | Strongest in direct matchups | LLM Stats |
| Cost-adjusted reasoning (GPQA + ARC-AGI-2 relative to price) | Gemini 3.1 Pro | Best price-to-score ratio of the three | LLM Stats |
Two gaps are worth naming instead of glossing over. Gemini 3.1 Pro does not have an independently confirmed GPQA Diamond or Intelligence Index score in the sources reviewed for this piece, even though it is frequently described in the same breath as the strongest cost-adjusted reasoning option. And GPT-5.6 Sol does not currently have a confirmed SWE-bench Pro or Terminal-Bench score in the trackers checked here, which matters if agentic coding reliability rather than general reasoning is the primary selection criterion.
Worth a mention for context: outside this specific three-way comparison, Claude Opus 4.8 remains Anthropic’s strongest model for agentic coding specifically, posting a 69.2 percent score on SWE-bench Pro, a harder successor to the original SWE-bench suite, according to Anthropic’s own disclosures corroborated by OpenRouter. Neither Sonnet 5, GPT-5.6 Sol, nor Gemini 3.1 Pro is being scored against that specific benchmark here, since the comparison is scoped to the three models most teams are actually choosing between for everyday, cost-sensitive traffic rather than the absolute best coding-agent money can buy.
Pricing Breakdown: What a Million Tokens Actually Costs
Per-token pricing is hard to reason about in the abstract, so the table below converts each model’s rate card into two realistic workloads: a short interactive session (100,000 input tokens, 20,000 output tokens) and a heavier batch job (1,000,000 input tokens, 200,000 output tokens).
| Model | Input / Output ($ per 1M tokens) | Small session (100K in / 20K out) | Batch job (1M in / 200K out) |
|---|---|---|---|
| Claude Sonnet 5 (est.) | ~$2.50 / ~$12.50 | $0.50 | $5.00 |
| GPT-5.6 Sol | $5.00 / $30.00 | $1.10 | $11.00 |
| Gemini 3.1 Pro (est.) | ~$2.00 / ~$12.00 | $0.44 | $4.40 |
The batch-job column is where the gap compounds fast. GPT-5.6 Sol costs more than double Sonnet 5’s estimated rate for the same volume, and it costs roughly 2.5 times Gemini 3.1 Pro’s estimated rate. Sonnet 5’s output-token price runs about 58 percent below Sol’s, a direct result of Anthropic pricing its mid-tier model at roughly half of Opus 4.8’s confirmed rate while Sol carries a premium over what GPT-5.5 charged before it.
The more surprising result sits between Sonnet 5 and Gemini 3.1 Pro. Anthropic markets Sonnet 5 as its value pick, but Gemini 3.1 Pro’s estimated rate still comes in slightly cheaper on both sides of the ledger, about 12 percent less on the small session and 12 percent less on the batch job. That gap is small enough that it should not be the deciding factor on its own, especially given Gemini’s pricing is a third-party estimate rather than a confirmed Google rate. But it undercuts the assumption that Sonnet 5 is automatically the budget option once Google’s preview model enters the conversation.
Two caveats apply across the board. None of these figures include prompt caching, batch-API discounts, or volume commitments, all of which can meaningfully cut effective cost on the two general-availability models. And Gemini’s numbers remain unconfirmed by Google directly, so a team building a long-term budget around that estimate should verify current pricing before signing off on a procurement decision.
Five Real-World Scenarios: Which Model Actually Gets Used
Benchmark tables answer “which model scores higher.” They do not answer “which model should my team actually deploy.” These five scenarios, drawn from the kinds of workloads that make up most production API traffic, illustrate how the tradeoffs above play out in practice.
- Customer support triage at scale. A SaaS company routing 50,000 support tickets a month toward auto-drafted responses does not need Opus-level reasoning on every ticket. Sonnet 5’s instruction-following lead and roughly half-of-Opus pricing make it the natural fit, since the job is mostly formatting, tone, and following a script reliably at volume.
- Long-document and video review. A legal-tech or media analytics team that needs to summarize hours of recorded calls, scan PDFs for clauses, or search a video archive gains the most from Gemini 3.1 Pro’s native audio and video input. Neither Sonnet 5 nor GPT-5.6 Sol accepts video natively, so that workload would otherwise require a separate transcription step before the model ever sees the content.
- An agentic coding assistant for a dev-tools startup. A team building a coding agent that plans multi-step changes across a repository is choosing primarily on reasoning depth and tool-use reliability, which is why GPT-5.6 Sol’s frontier GPQA Diamond score and Terminal-Bench-style agentic focus make it a reasonable anchor model to evaluate first, even at the premium price.
- High-volume marketing copy generation. An e-commerce brand generating thousands of product descriptions and ad variants a week cares more about writing quality and brand-voice consistency than raw IQ. Sonnet 5’s category lead in writing quality directly targets that job, and the lower per-token cost matters more when the same prompt structure runs thousands of times a day.
- A cost-constrained startup building its first AI feature. A two-person startup adding a chat feature to its product before any real usage data exists benefits from starting with whichever model has the lowest realistic per-call cost while it validates demand, which points toward Sonnet 5 or Gemini 3.1 Pro over Sol’s higher output rate, with the caveat that Gemini’s preview status carries its own migration risk down the line.
The pattern across all five scenarios is the same: the “best” model depends entirely on which specific capability the job actually needs, not which model posted the highest score on the benchmark a team happened to read last.
Multimodal Reach and Context Window Compared
If a workload touches anything beyond plain text, the gap between these three widens considerably. Gemini 3.1 Pro is the clear leader here. Google’s own model card lists native support for text, audio, images, video, PDFs, and entire code repositories inside a single context window, a breadth neither competitor currently documents at the same level. That makes it the obvious starting point for jobs like transcribing and summarizing a video call, extracting structured data from scanned PDFs, or feeding an entire multi-file codebase into one prompt for analysis.
Claude Sonnet 5 documents text, image, and document input, inherited from the same family architecture as Opus 4.8, which covers most business use cases, contracts, screenshots, spreadsheets, without stretching into audio or video. GPT-5.6 Sol’s officially documented input modalities are narrower still: text and image at the Sol tier, though OpenAI has signaled broader modality support may arrive across the GPT-5.6 family over time.
Context windows across the three now cluster close together, a genuine shift from a year earlier when window size was a clear differentiator. GPT-5.6 Sol discloses the largest figure at 1,050,000 tokens, Gemini 3.1 Pro sits at 1,000,000, and Sonnet 5’s window is not separately disclosed but presumably tracks Opus 4.8’s 1,000,000-token architecture. Max output tokens tell a different story: Sol allows up to 128,000 tokens per response, more than double Gemini 3.1 Pro’s 64,000-token ceiling, which matters for jobs that generate long documents in a single pass rather than jobs that mostly read long input and return a short answer.
The practical effect of the multimodal gap shows up in preprocessing costs that never appear on a pricing page. A team feeding video into Sonnet 5 or GPT-5.6 Sol today needs a separate transcription or frame-extraction step before either model ever sees the content, adding latency, an extra vendor dependency, and its own per-minute cost. Route that same video straight into Gemini 3.1 Pro and the preprocessing step disappears. Whether that tradeoff is worth Gemini’s preview-status risk depends on how much of a given workload is genuinely multimodal versus text with the occasional attached image.
Agentic Coding and Tool Use
All three vendors support function calling and tool use, but the API shapes differ enough that switching between them is not a drop-in change. The snippet below illustrates the structural difference in how a request is built, not production-ready code.
# Anthropic (Claude Sonnet 5)
response = client.messages.create(
model="claude-sonnet-5",
max_tokens=4096,
tools=[tool_schema],
messages=[{"role": "user", "content": prompt}]
)
# OpenAI (GPT-5.6 Sol)
response = client.responses.create(
model="gpt-5.6-sol",
input=prompt,
tools=[tool_schema]
)
# Google (Gemini 3.1 Pro)
response = client.models.generate_content(
model="gemini-3.1-pro-preview",
contents=prompt,
tools=[tool_schema]
)
Beyond syntax, the practical difference shows up in longer agentic runs rather than single-turn calls. A model that writes correct code in isolation but forgets earlier context by step fifteen is not actually useful for autonomous work, which is why Anthropic’s Opus-tier Terminal-Bench results get cited so often in agentic-coding discussions even though Sonnet 5 itself is not primarily marketed on that axis. GPT-5.6 Sol’s strength here leans on its GPQA Diamond lead translating into more reliable multi-step reasoning during a long tool-use chain, according to LLM Stats’ coverage of the release, though it lacks an independently confirmed long-horizon agentic score in the same tracker.
Gemini 3.1 Pro’s agentic strength is less about raw reasoning depth and more about not needing a separate ingestion pipeline before the agent starts working. An agent that has to review a video walkthrough, a PDF spec, and a codebase in the same task benefits from Gemini’s single-context multimodal input in a way neither Sonnet 5 nor GPT-5.6 Sol can currently match without extra tooling bolted on.
Which Model Fits Which Job: Six Use-Case Recommendations
- Choose Claude Sonnet 5 for high-volume, text-heavy production traffic where instruction-following and consistent brand voice matter more than frontier reasoning: support automation, internal documentation, and content generation at scale.
- Choose GPT-5.6 Sol when a task genuinely needs the strongest available reasoning score in this comparison and the budget supports its output-token premium: complex research assistants, technical Q&A products, and reasoning-heavy agents where accuracy failures are expensive.
- Choose Gemini 3.1 Pro for anything genuinely multimodal: video review, large PDF extraction, or feeding an entire codebase into one prompt, provided the team can tolerate building on a model still labeled Public Preview.
- Choose Sonnet 5 over Sol when the workload is cost-sensitive and text-only, since the roughly 58 percent lower output-token price compounds fast at scale for a two-point difference on a composite reasoning index.
- Choose Sol over Gemini 3.1 Pro when general availability and a confirmed release status matter more than price, since Gemini’s preview label carries real repricing and rate-limit risk that a production contract may not tolerate.
- Avoid defaulting to any single flagship tier for mixed workloads. Teams running both routine and hard tasks through the same model are very likely overpaying on the routine share. Routing cheap, high-volume calls to Sonnet 5 or a comparable budget tier and reserving Sol or Opus 4.8 for the genuinely hard steps is the pattern all three labs now build their own pricing around.
Migration Guide: Moving Between Sonnet 5, GPT-5.6, and Gemini 3.1 Pro
Switching a production integration from one of these models to another is rarely a one-line change, even when the underlying task stays the same. The steps below cover the parts of a migration that most commonly break.
- Audit modality use first. Before touching any code, confirm which input types the current integration actually sends. A migration from Gemini 3.1 Pro to either Sonnet 5 or GPT-5.6 Sol will fail immediately on any request that includes audio or video, since neither accepts those inputs natively.
- Remap the client SDK and request shape. Anthropic’s messages.create, OpenAI’s responses.create, and Google’s generate_content calls use different parameter names for the same concepts, model, prompt, and tool definitions. Budget time for a thin adapter layer rather than assuming a find-and-replace will work.
- Re-test tool-calling schemas independently. Function-calling syntax is not identical across the three, and a tool schema that validates cleanly on one API can silently fail or get ignored on another. Run the full tool-use test suite again rather than trusting that a passing test on the old model implies a passing test on the new one.
- Rebuild prompt-caching and cost assumptions. Cached-input discounts and batch-API pricing are not identical across vendors, and a cost model built around one provider’s caching behavior can understate real spend on another. Recalculate the pricing table above against actual traffic volume before rolling out broadly.
- Stage the rollout behind a feature flag. Run the new model against a percentage of live traffic while keeping the old integration as a fallback, and compare output quality on the writing-style and instruction-following dimensions specifically, since those are the categories most likely to shift in ways a pure accuracy metric will not catch.
- Plan for Gemini’s preview status specifically. Any migration onto Gemini 3.1 Pro should include a rollback plan, since Google has not committed to a stable release date and preview pricing or rate limits could change before general availability.
Pros and Cons of Each Model
Claude Sonnet 5
- Pro: leads writing quality and instruction-following among the models tracked at its price point
- Pro: priced at roughly half of Opus 4.8’s confirmed rate, the cheapest general-availability option of the three on output tokens
- Pro: general availability since June 30, 2026, with no preview-status risk
- Con: trails both Opus 4.8 and GPT-5.6 Sol on composite reasoning scores
- Con: narrower multimodal input than Gemini 3.1 Pro, no native audio or video
- Con: context window and max output figures are not separately disclosed, making capacity planning harder
GPT-5.6 Sol
- Pro: leads this comparison on GPQA Diamond at 94.6 percent and on the Artificial Analysis Intelligence Index
- Pro: largest disclosed context window of the three at 1,050,000 tokens, with the highest max output ceiling at 128,000 tokens
- Pro: most recent knowledge cutoff of the three, February 16, 2026
- Con: most expensive of the three on output tokens at $30.00 per million, more than double Sonnet 5’s estimated rate
- Con: narrowest officially documented multimodal input, limited to text and image at the Sol tier
- Con: lacks an independently confirmed SWE-bench Pro or Terminal-Bench score in the trackers reviewed here
Gemini 3.1 Pro
- Pro: broadest native multimodal input of the three, spanning text, audio, image, video, PDF, and code repositories
- Pro: estimated pricing slightly undercuts even Sonnet 5 on both input and output tokens
- Pro: frequently cited as the strongest cost-adjusted reasoning option once price is weighed against GPQA Diamond and ARC-AGI-2 results
- Con: still labeled Public Preview more than five months after its February 2026 debut, with no confirmed general-availability date
- Con: pricing is a third-party estimate rather than an official Google rate, creating budget uncertainty
- Con: lowest max output ceiling of the three at 64,000 tokens
Where These Three Sit Against the Rest of the Field
None of these three models exist in a vacuum, and it is worth zooming out before settling on a verdict. Anthropic’s own Claude Opus 4.8 remains the mainline flagship above Sonnet 5, leading on agentic coding with a 69.2 percent SWE-bench Pro score, and Anthropic’s newer Mythos-class Claude Fable 5 sits above Opus 4.8 again on raw benchmark scores at a different price and availability tier. Our Claude Fable 5 vs Opus 4.8 breakdown covers that higher tier in detail, and our dedicated Opus 4.8 vs GPT-5.5 vs Gemini 3.1 Pro comparison and Opus 4.8 vs GPT-5.6 vs Gemini 3.1 Pro comparison both cover the full-flagship matchup this piece deliberately does not repeat.
GPT-5.5, the model Sol replaced as ChatGPT’s default, is still reachable via API for teams that have not migrated, and xAI’s Grok 4.3 is frequently cited alongside this trio in third-party trackers, with pricing reported around $1.25 per million input tokens and $2.50 per million output tokens, undercutting all three models covered here on paper, though xAI has not confirmed that figure on its own site as clearly as the other vendors have. Our Grok 4.3 vs Gemini 3.1 Pro breakdown digs into that specific matchup.
On the cost-efficient end, open-weight models such as DeepSeek and Kimi K2.6 get cited as delivering much of the capability of these mid-tier and flagship models at a fraction of the price, a pattern our best AI model for coding comparison covers in detail for teams whose primary constraint is budget rather than benchmark leadership. The practical takeaway is that Sonnet 5, GPT-5.6 Sol, and Gemini 3.1 Pro represent a specific slice of the market, the everyday tier most teams actually build against, not the only tier available and not the cheapest one either.
The Verdict: Which Model Should You Actually Use in 2026
There is no single winner here, and the data does not support pretending otherwise. GPT-5.6 Sol wins on raw reasoning, topping GPQA Diamond at 94.6 percent and the Intelligence Index at 59 in its max configuration, and it earns that lead honestly with the most recent knowledge cutoff and the largest context window of the three. It also costs the most by a wide margin, more than double Sonnet 5’s estimated output rate and roughly 2.5 times Gemini 3.1 Pro’s.
Claude Sonnet 5 wins on the metric that actually drives most production budgets: cost per reliable output. A two-point gap on a composite reasoning index does not justify double the price for a support bot, a copy generator, or an internal tool that mostly needs to follow instructions correctly at volume. For that category of workload, which covers the bulk of real API traffic according to Anthropic’s own framing of the release, Sonnet 5 is the more rational default.
Gemini 3.1 Pro wins the category neither of the other two can touch: genuine multimodal input in a single context window, at a price that undercuts even Anthropic’s value tier. The catch is real, though. A five-month-old preview label with no confirmed release date is a legitimate risk for any team building a long-term product on top of it, and the pricing is still an estimate rather than a locked Google rate.
The honest verdict: match the model to the job rather than picking a permanent winner. Default routine, high-volume, text-based traffic to Sonnet 5. Reserve GPT-5.6 Sol for tasks where the reasoning premium is worth paying for. Bring in Gemini 3.1 Pro specifically when a workload is genuinely multimodal and the team can tolerate preview-status risk. Teams that pick one model and route every request through it, regardless of the job, are very likely leaving either quality or budget on the table.
Frequently Asked Questions
Is Claude Sonnet 5 as capable as Claude Opus 4.8?
Not on composite reasoning. Sonnet 5 scored 57 on the Artificial Analysis Intelligence Index against Opus 4.8’s 61 in the same comparative tracking. Sonnet 5 does lead its own family on writing quality and instruction-following, and it costs roughly half of what Opus 4.8 charges per million tokens, so the gap matters less for workloads that do not need Opus-tier reasoning.
Is GPT-5.6 Sol worth the price premium over Sonnet 5?
It depends on the job. Sol’s 94.6 percent GPQA Diamond score and higher Intelligence Index justify the cost for reasoning-heavy tasks, but its output-token price runs more than double Sonnet 5’s estimated rate for a two-point gap on the composite index, which is a hard premium to justify for routine, high-volume traffic.
Can I use Gemini 3.1 Pro in a production app if it is still in preview?
Teams do it, but it carries real risk. Google has kept Gemini 3.1 Pro in Public Preview since February 2026 with no confirmed general-availability date, and its pricing is a third-party estimate rather than an official Google rate. Plan for possible repricing or rate-limit changes before committing a long-term product to it.
Which of the three models is cheapest for high-volume traffic?
On estimated per-token rates, Gemini 3.1 Pro is slightly cheaper than Sonnet 5 on both input and output tokens, with GPT-5.6 Sol the most expensive by a wide margin. Since Gemini’s rate is an unconfirmed estimate and Sonnet 5’s is a general-availability model with a stable release status, the practical answer depends on how much pricing certainty a team needs.
Which model has the strongest coding performance?
None of the three is benchmarked here on SWE-bench Pro or Terminal-Bench, the two agentic-coding benchmarks Anthropic’s own Opus 4.8 leads at 69.2 percent and 74.2 percent respectively. Among this specific trio, Gemini 3.1 Pro is cited as strongest in direct coding-arena matchups, though that is a relative ranking rather than a percentage score.
What is the real difference in context window size?
Smaller than it used to be. GPT-5.6 Sol discloses 1,050,000 tokens, Gemini 3.1 Pro sits at 1,000,000, and Sonnet 5’s figure is not separately published but likely tracks the same 1,000,000-token architecture Opus 4.8 uses. The bigger practical difference is max output: Sol allows up to 128,000 tokens per response against Gemini’s 64,000-token ceiling.
How hard is it to migrate an app between these three APIs?
Harder than a find-and-replace. Each vendor uses a different SDK method and parameter structure for the same core actions, and tool-calling schemas need independent re-testing since a passing test on one API does not guarantee the same behavior on another. See the migration guide above for the full checklist.
Do any of these three offer a free tier for testing?
Each vendor offers some form of limited free or trial access through its respective consumer app or developer console, but the pricing figures in this comparison refer specifically to production API rates per million tokens, not consumer subscription plans. Confirm current trial terms directly with each vendor before assuming free-tier access covers a specific workload.
Related Coverage
- Claude Sonnet 5 Debuts: 57 Score, Half the API Cost
- Opus 4.8 vs GPT-5.5 vs Gemini 3.1: 8-Point SWE Gap
- Opus 4.8 vs GPT-5.6 vs Gemini 3.1 Pro: $18 Price Gap
- Claude Fable 5 vs Opus 4.8: 11-Point SWE-Bench Gap
- Best AI Model for Coding: DeepSeek Costs 20x Less
- Grok 4.3 vs Gemini 3.1 Pro: 4x Output Price Gap
For the wider picture of where every major model launched in 2026 stacks up, see our best AI models 2026 hub. Sources consulted for the figures in this piece include Anthropic’s official announcements, Google DeepMind’s model documentation, Artificial Analysis’ independent benchmark tracker, OpenRouter’s aggregated model and pricing listings, the SWE-bench benchmark suite, and LLM Stats.


