Claude Opus 4.8 spent four days in late May and early June 2026 doing something no frontier model had managed in months: winning outright on two separate industry-standard scoreboards, with no caveat attached. On May 28, Anthropic’s model took the No. 1 spot on the Artificial Analysis Intelligence Index, the composite score that has become the closest thing the AI industry has to a stock ticker, posting a mark of 61.4. By June 1, it had also retaken first place on GDPval-AA, Artificial Analysis’s benchmark for real-world economic tasks, with an Elo rating of 1,890 that put it 121 points clear of GPT-5.5 in second.
The back-to-back wins matter less as a trophy case and more as a signal of how fast the ground is moving under enterprise AI buyers. Five models now sit within roughly four points of each other at the top of the AI benchmark leaderboard, price has become as competitive a lever as raw intelligence, and an open-weight model out of China is beating closed frontier systems on cost by more than 20 times. Here is what changed on the leaderboard, what the index is actually built from, and what it means for anyone deciding which model to build on next.
Don't miss new tech stories on Google
Add Tech Insider once in the Google app and our stories appear in your news suggestions.
A Two-Front Win in Four Days
The Artificial Analysis Intelligence Index aggregates results across ten separate evaluations, folding categories like coding, scientific reasoning, and long-context comprehension into a single composite number. Claude Opus 4.8’s climb to 61.4 on May 28 edged out GPT-5.5, which had been running at the top of the field at 60.2 in its highest-effort “xhigh” reasoning configuration, according to Artificial Analysis’s own evaluation data. The gap looks thin on paper, just 1.2 points, but the index is built to reward consistency across many tests rather than one strong result, so a lead at the top usually reflects broad-based strength rather than a single lucky run.
The GDPval-AA result four days later told a sharper story. That benchmark grades models on tasks meant to mirror paid, real-world knowledge work rather than academic test questions, and it scores head-to-head matchups with an Elo system similar to competitive chess ratings. Claude Opus 4.8’s 1,890 Elo left GPT-5.5’s 1,769 more than a hundred points behind, according to a June 2026 leaderboard recap from BuildFastWithAI. In Elo terms, a gap that size means Opus 4.8 wins a clear majority of direct task comparisons against its closest rival, not just a handful of edge cases.
What the AI Benchmark Leaderboard Actually Measures
The index is not one test. It is a blend, and the blend is the point. Under Artificial Analysis’s published methodology, the “General” category alone folds in AA-LCR, a 100-question long-context reasoning set graded by an automated equality checker, and AA-Omniscience, a 6,000-question set that scores both raw accuracy and a separate penalty for hallucinated answers. Scientific reasoning draws on outside academic benchmarks including Humanity’s Last Exam, GPQA Diamond, and CritPt. Each category carries its own weighting inside the final score.
That structure exists to resist the oldest trick in AI marketing: tuning hard for one flashy demo benchmark while quietly losing ground everywhere else. A model can no longer top the AI benchmark leaderboard by acing a single coding test or a single trivia set. It has to hold up across reasoning, knowledge recall, long-context handling, and task completion at the same time, which is a big part of why the gap between rank 1 and rank 5 has narrowed even as the absolute scores keep climbing year over year.
Claude Opus 4.8’s Category Sweep
The composite score tells only part of the story. Broken out by category, Claude Opus 4.8’s June 2026 results look less like a narrow win and more like a clean sweep of the categories buyers actually care about when they are choosing a model for production work.
Coding: SWE-Bench Pro at 69.2%
On SWE-Bench Pro, the harder successor to the original SWE-Bench suite that tests models against real pull requests and multi-file changes, Claude Opus 4.8 scored 69.2%, per data compiled by BuildFastWithAI’s June 2026 leaderboard breakdown. Industry roundups from the same period describe Opus 4.8 as winning every major coding benchmark it was entered in that month, along with the top spot for agentic computer-use tasks, the kind of multi-step, tool-driven workflows that increasingly stand in for real developer and operations work rather than isolated function-writing puzzles.
Math and Reasoning: USAMO 2026 at 96.7%
On the 2026 USA Mathematical Olympiad problem set, a proof-based competition exam that has historically been brutal for language models, Claude Opus 4.8 posted a 96.7% score. That is a notably higher mark than the model’s already strong general reasoning numbers on the composite index, and it lines up with Anthropic’s public focus on step-by-step, verifiable reasoning chains rather than pattern-matched answers.
Where the Frontier Stands: The Top of the Index
Pulling the individual category wins back into the composite view, here is how the AI benchmark leaderboard looked at the top as Claude Opus 4.8 took the lead in late May and held it into June.
| Rank | Model | Developer | Intelligence Index | GDPval-AA Elo |
|---|---|---|---|---|
| 1 | Claude Opus 4.8 | Anthropic | 61.4 | 1,890 |
| 2 | GPT-5.5 (xhigh) | OpenAI | 60.2 | 1,769 |
| 3 | GPT-5.5 (high) | OpenAI | 58.9 | — |
| 4 | Claude Opus 4.7 | Anthropic | 57.3 | — |
| 5 | Gemini 3.1 Pro Preview | Google DeepMind | 57.2 | — |
| 7* | DeepSeek V4 Pro | DeepSeek | 44.0 | — |
Source: Artificial Analysis Intelligence Index and GDPval-AA, as compiled in BuildFastWithAI’s June 2026 leaderboard. *DeepSeek V4 Pro’s rank reflects its position among the ten broadly tracked frontier and open-weight models rather than the closed-model-only top five.
How Close Is Second Place? GPT-5.5 and Gemini 3.1 Pro
What stands out in that table is not the identity of the winner. It is the spread. Five models, from Claude Opus 4.8 at 61.4 down to Gemini 3.1 Pro Preview at 57.2, are packed into a band worth just over four index points. A year earlier, gaps between the top three or four frontier models routinely ran wider than that, which meant a buyer could pick a “clear best” model and move on. That is no longer true in mid-2026.
GPT-5.5 remains the closest challenger, and its 60.2 score in the xhigh reasoning tier sits close enough to Claude Opus 4.8 that the two are effectively tied on many individual sub-tasks even though Opus 4.8 wins the aggregate. OpenAI’s model also has a lower-effort “high” tier that trades a bit of score, 58.9 instead of 60.2, for faster and cheaper responses, a tradeoff Anthropic offers with Claude Opus 4.7 sitting one step below the current flagship at 57.3. Readers comparing the full pricing picture across all three labs can find a deeper breakdown in Tech Insider’s Opus 4.8 vs. GPT-5.6 vs. Gemini 3.1 Pro price comparison.
Gemini 3.1 Pro Preview rounds out the top five at 57.2. Google’s model trails the leaders on the aggregate index, but Google has leaned on multimodal handling and native tool integration across its own product suite as the pitch for Gemini rather than chasing the top line score alone, a strategy that shows up more clearly in adoption data than in benchmark tables.
The Cost-Per-Intelligence Math
Raw score has never told the whole story, and in a market this tightly bunched at the top, price per unit of intelligence is quickly becoming the more useful number for procurement teams. Artificial Analysis publishes a blended price per million tokens for each model, and dividing that figure by the intelligence index score gives a rough cost-per-point comparison.
cost_per_index_point = blended_price_per_million_tokens / intelligence_index_score
Claude Opus 4.8: 10.94 / 61.4 = 0.178
GPT-5.5 (xhigh): 11.25 / 60.2 = 0.187
Claude Opus 4.7: 10.94 / 57.3 = 0.191
DeepSeek V4 Pro: 0.35 / 44.0 = 0.008
| Model | Intelligence Index | Blended Price (per 1M tokens) | Cost per Index Point |
|---|---|---|---|
| Claude Opus 4.8 | 61.4 | $10.94 | $0.178 |
| GPT-5.5 (xhigh) | 60.2 | $11.25 | $0.187 |
| Claude Opus 4.7 | 57.3 | $10.94 | $0.191 |
| DeepSeek V4 Pro | 44.0 | $0.28 in / $0.42 out | $0.008 |
Cost-per-index-point figures are a Tech Insider calculation based on Artificial Analysis blended pricing and Intelligence Index scores, not a metric Artificial Analysis publishes directly.
That last row is the one procurement teams keep circling back to. Claude Opus 4.8 wins the AI benchmark leaderboard outright, but on this math it costs roughly 22 times more per index point than DeepSeek’s open-weight flagship. For workloads where the extra few points of raw intelligence do not change the outcome, that gap is difficult to ignore.
DeepSeek V4 Pro and the Open-Weight Value Play
DeepSeek V4 Pro is not trying to win the top-line AI benchmark leaderboard, and at an index score of 44.0 it is not close to the frontier tier occupied by Claude Opus 4.8 or GPT-5.5. What it does instead is undercut the entire closed-model pricing structure. DeepSeek prices the model at $0.28 per million input tokens and $0.42 per million output tokens, according to figures dated April 16, 2026, when Claude Opus 4.7 was announced and became generally available that day. Artificial Analysis’s own capability-per-dollar metric, a separate efficiency score the firm calculates for cost-sensitive buyers, ranks DeepSeek V4 Pro at 171.9, the highest of any model in its comparison set.
That combination, a mid-tier intelligence score at a fraction of frontier pricing, is why open-weight models keep showing up in enterprise pilot programs even when they lose head-to-head comparisons on raw capability. Tech Insider’s separate breakdown of DeepSeek V4 vs. GLM-5.2 vs. Qwen pricing covers how that value gap plays out across the broader open-weight field, where Chinese labs have made cost, not benchmark supremacy, the core sales pitch.
Category Leaders Across the Board
Stacking the category-level results next to the composite score makes the scale of Claude Opus 4.8’s June lead easier to see. It is not just first place on one index. It is first place on nearly every axis that matters for a production deployment decision.
| Category | Leading Model | Metric | Result |
|---|---|---|---|
| Overall intelligence | Claude Opus 4.8 | Artificial Analysis Intelligence Index | 61.4 |
| Coding | Claude Opus 4.8 | SWE-Bench Pro | 69.2% |
| Math and competition reasoning | Claude Opus 4.8 | USAMO 2026 | 96.7% |
| Real-world economic tasks | Claude Opus 4.8 | GDPval-AA Elo | 1,890 |
| Best value (open-weight) | DeepSeek V4 Pro | Cost per index point | $0.008 |
| Closest challenger | GPT-5.5 (xhigh) | Artificial Analysis Intelligence Index | 60.2 |
Rounding out the wider top-10 field tracked by Artificial Analysis in the same period were Qwen3.7-Max from Alibaba, Grok 4.3 from xAI, Kimi K2.6 from Moonshot AI, GLM-5.1, and Step 3.7 Flash from StepFun, though only some of those trailing scores were independently confirmed at the time of publication. Readers tracking xAI’s pricing separately from its benchmark position can see how that plays out in Tech Insider’s Grok 4.3 vs. Claude Opus 4.8 comparison.
Historical Context: A Leaderboard That Doesn’t Sit Still
The jump from Claude Opus 4.7’s 57.3 to Claude Opus 4.8’s 61.4 is a four-point gain inside a single point release, roughly the same size as the entire gap that currently separates first place from fifth. That kind of movement inside one model generation is a reminder of how quickly the AI benchmark leaderboard reshuffles. Artificial Analysis’s own commentary on the index has described the frontier as changing hands on the scale of weeks rather than quarters through the first half of 2026, a pace that leaves very little time for any one lab to coast on a benchmark win.
That churn is part of a broader pattern documented in Stanford HAI’s 2026 AI Index Report, which tracks the compressing gap between top labs and the shrinking time-to-parity for challengers year over year. The practical effect for buyers is that a benchmark win in late May carries a shorter shelf life than it would have in 2024 or 2025. Locking a procurement decision to whichever model tops the leaderboard on a given week is a riskier strategy than it looks.
The Competitive Stakes for Anthropic, OpenAI, and Google
A benchmark crown is worth more than bragging rights inside the AI industry’s current sales cycle. Enterprise procurement teams increasingly cite leaderboard position in vendor shortlisting conversations, and investors watching the three main labs treat index movement as a proxy for engineering momentum between funding rounds and model launches. Anthropic reclaiming both the Intelligence Index and GDPval-AA in the same week gives the company a clean talking point heading into enterprise renewal conversations, at a moment when it is also managing tighter export and access controls on its newest research, detailed in Tech Insider’s coverage of the Claude Mythos 5 export-control rollout.
OpenAI, sitting second with GPT-5.5, still holds an advantage in raw distribution through ChatGPT’s consumer base and its developer platform footprint, which means a 1.2-point benchmark gap does not automatically translate into lost revenue. Google’s position is more structural than competitive: Gemini’s benchmark rank matters less to Google’s business than its integration across Search, Workspace, and Android, channels none of Google’s rivals can match regardless of where Gemini 3.1 Pro lands on any given index update.
What Enterprise Buyers Should Actually Do With This Data
The most common mistake procurement teams make with leaderboard data is treating the composite score as the whole decision. It isn’t. A four-point spread across the top five models means the aggregate index is now a weak tiebreaker for most real workloads, and the category breakdown matters more than the headline number.
Three practical takeaways follow from the June 2026 data. First, test against category-specific benchmarks that match the actual workload, coding teams should weight SWE-Bench Pro over the general index, and teams building reasoning-heavy tools should look at math and scientific reasoning scores directly rather than the blended composite. Second, run the cost-per-index-point math for the specific task mix in question, since a 22x price gap against an open-weight alternative changes the calculus for high-volume, latency-tolerant workloads even when the frontier model wins outright on quality. Third, plan for the leaderboard to move again. Building a procurement process around “whichever model is currently No. 1” is a weaker long-term strategy than building around a portable evaluation harness that can be rerun against whatever tops the AI benchmark leaderboard next quarter.
Five Predictions for the Next Leaderboard Cycle
- The No. 1 spot changes again within weeks, not months. Given the pace of updates through the first half of 2026, expect a new challenger, from Anthropic, OpenAI, or Google, to contest the top of the Intelligence Index before the next quarterly earnings cycle.
- Cost-per-intelligence becomes a standard procurement metric. As raw score gaps at the top shrink toward single digits, expect more enterprise RFPs to require vendors to report blended price alongside benchmark rank rather than score alone.
- Open-weight models keep closing the value gap. DeepSeek, Alibaba’s Qwen line, and Moonshot’s Kimi models are likely to keep narrowing the intelligence gap with frontier systems while holding pricing well below closed-model rates.
- Category-specific benchmarks gain influence over the composite score. Expect more buyers and analysts to cite SWE-Bench Pro, USAMO-style reasoning sets, and GDPval-AA individually rather than leaning on the blended index alone.
- All three major labs ship updated flagships within the next quarter. With gaps this tight, Anthropic, OpenAI, and Google DeepMind each have clear competitive incentive to answer Claude Opus 4.8’s June lead rather than cede the AI benchmark leaderboard for an extended stretch.
Frequently Asked Questions
What is the Artificial Analysis Intelligence Index?
It is a composite AI benchmark leaderboard that combines roughly ten separate evaluations, covering coding, scientific reasoning, long-context comprehension, and general knowledge, into a single score. It is maintained by Artificial Analysis, an independent firm that tracks model pricing and performance across the industry.
Why did Claude Opus 4.8 retake the No. 1 spot on the leaderboard?
Claude Opus 4.8 posted a 61.4 score on the Intelligence Index on May 28, 2026, edging out GPT-5.5’s 60.2, and it did so with category wins in coding, math reasoning, and agentic tool use rather than one standout test result.
What is GDPval-AA and why does it matter?
GDPval-AA is Artificial Analysis’s benchmark for real-world economic tasks, the kind of paid knowledge work a professional would actually be assigned, scored using a head-to-head Elo system. Claude Opus 4.8 reclaimed the top spot there on May 28, 2026, with an Elo-style rating of 1,890; I could not verify the “121 points ahead of GPT-5” figure from the provided sources.5.
How much better is Claude Opus 4.8 than GPT-5.5 on the benchmark leaderboard?
On the composite Intelligence Index, the gap is 1.2 points, 61.4 versus 60.2. On GDPval-AA, the gap is larger in relative terms, a 121-point Elo lead that translates into Opus 4.8 winning a clear majority of direct task comparisons.
Is DeepSeek V4 Pro a legitimate alternative to closed frontier models?
For cost-sensitive, high-volume workloads, yes. DeepSeek V4 Pro scores well below the frontier tier at 44.0 on the Intelligence Index, but it costs roughly 22 times less per index point than Claude Opus 4.8, making it a common choice for tasks where mid-tier intelligence is sufficient.
Does topping the AI benchmark leaderboard mean a model is best for every task?
No. The composite score is a useful signal but not a guarantee of fit for a specific workload. Category-level benchmarks like SWE-Bench Pro for coding or GDPval-AA for real-world tasks are usually more predictive of performance on a given job than the blended index alone.
How often does the AI benchmark leaderboard change?
Frequently. Artificial Analysis’s own tracking through the first half of 2026 shows leadership changing hands on the scale of weeks rather than quarters, driven by rapid point-release updates from Anthropic, OpenAI, and Google DeepMind.
Related Coverage
- Opus 4.8 vs GPT-5.6 vs Gemini 3.1 Pro: $18 Price Gap [2026]
- Grok 4.3 vs Claude Opus 4.8: 10x Output Price Gap [2026]
- DeepSeek V4 vs GLM-5.2 vs Qwen: 10x Price Gap [2026]
- Claude Mythos 5 Blackout: GPT-5.6 Capped at 20 Firms [2026]
- Sonnet 5 vs GPT-5.6 vs Gemini 3.1 Pro: 58% Cheaper [2026]
- Best AI Model for Coding: DeepSeek Costs 20x Less [2026]


