SYSTEM NOTICE

Auto translation by AI. Be sure, accuracy, nuances and authorial intent may not be fully reflected.
見出し画像

The Age of Choosing AI in Three Tiers — How 5% Became 31%

Foresight Lab.

A quick note: note.com recently added AI auto-translation, so putting out a separate English edition may well be redundant. Even so, I wanted to try a hand-written English version as an experiment. This is the English edition of a piece written for Japanese business readers; if you reference it, a mention of Foresight Lab. is appreciated.

——————————————————
The short version
——————————————————

OpenAI's new GPT-5.6 family isn't a single model — it comes in three tiers: Sol, Terra, and Luna.
① The top-end Sol is available only to select partners, at the request of the U.S. government.
② On a life-sciences reasoning benchmark, the pass rate jumped from under 5% to 31.5%.
③ Terra delivers roughly the previous generation's performance at about half the price, and Luna goes lower still.
Together these signal that AI's competition is shifting from "use the single smartest model" to "match the model to the job."

——————————————————
The full piece
——————————————————

On a reasoning benchmark where the previous generation scored under 5%, the newest model reached 31.5%.

And that newest model is doing something unprecedented: undergoing a U.S. government safety review before general release. What stopped me wasn't the raw pace of AI progress, but how that progress rebounds onto business decisions.

In late June, OpenAI announced a new model family, GPT-5.6. What's distinctive is that it isn't one model but three tiers — Sol, Terra, and Luna. Sol is the top-end brain; Terra is the price-performance balance; Luna is the low-cost option for high-volume work. That three-tier structure itself is the business shift I found most notable.

The first point is that Sol's release is restricted at the request of the U.S. government.

Sol far outperforms prior models in coding and cybersecurity, and posts strong results on benchmarks used for offensive security testing while using fewer tokens. Because of that capability, U.S. authorities asked that it be offered first only to a small group of trusted partners until a formal evaluation framework is in place. In other words, the very scope of a frontier model's release is now a matter of policy, not just technical assessment. This isn't a distant concern for Japanese firms either — we are entering an era where access to the most advanced models may hinge on your contract terms and compliance posture.

The second point is a new benchmark that measures judgment itself, called GeneBench-Pro.

It consists of 129 life-sciences problems and tests the ability to choose the right analytical method from messy, unorganized experimental data, revising hypotheses along the way to reach an answer. Rather than textbook one-shot Q&A, it recreates the chain of judgment calls a working researcher faces daily. Where a human expert might spend 20 to 40 hours on such a task, AI can attempt it for a few dollars of compute. The top model reaching 31.5% at the highest difficulty setting also means, conversely, that it still errs on nearly 70% of problems. Neither overestimation nor underestimation is wise; the number honestly tells us that automating judgment-heavy work is still a work in progress.

The third point is the price disruption the three tiers bring.

Terra holds roughly the same performance as the previous top model at about half the cost. Luna goes lower still for routine tasks like summarization and classification. Many companies have operated on "just use the smartest model available." But this tiering of price signals a stage where the design itself — judging the accuracy and cost each task requires, and deploying multiple models accordingly — becomes a competitive difference. Top-tier models for high-stakes, no-failure work; low-cost models for everyday bulk processing. How well you route between them will heavily shape your AI costs going forward.

Here's the one thing I'd like readers to take away. Take an inventory of the work you routinely hand to AI, and sort each task by whether it truly needs top-tier judgment or can be treated as routine processing. In most cases, the bulk of the work can be handled well by low-cost models, and the situations that genuinely demand near-human judgment should narrow to a select few. As much as keeping up with AI's progress, rethinking how you choose the model you already use looks set to be where the difference is made in the coming months.

Foresight Lab.

━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

Foresight Lab. publishes on AI and technology. If this was useful, a follow helps you catch what's next.


いいなと思ったら応援しよう!