Fugu Ultra vs Fable 5 vs Mythos: How the Sakana Orchestrator Stacks Up Against the Frontier

Fugu Ultra vs Fable 5 vs Mythos, honestly compared: Sakana's orchestrator claims parity, but it calls frontier models. See what that means.

Ashley Innocent

Ashley Innocent

22 June 2026

Fugu Ultra vs Fable 5 vs Mythos: How the Sakana Orchestrator Stacks Up Against the Frontier

Apidog for Enterprise

On-Premises Deploy

SSO & RBAC

SOC 2 Compliant

Explore Apidog Enterprise

Fugu Ultra is Sakana AI’s top variant of Fugu, and the release frames it as a peer to the current frontier, not a conqueror of it. Per Sakana, Fugu Ultra “stands shoulder-to-shoulder with leading models like Fable 5 and Mythos Preview” on engineering, scientific, and reasoning benchmarks. That is a parity claim, not a “beats” claim. The important catch: Fugu is an orchestrator that calls other vendors’ models, so it competes in a different category than the single Anthropic models it sits beside. The full details live on the Sakana Fugu release page, and you can read our deeper breakdown in what is Sakana Fugu.

button

What you are comparing

Fugu is a multi-agent orchestration system presented as one foundation model behind one OpenAI-compatible API. Sakana describes it as a trained language model specialized in delegation, agent communication, and work synthesis. It coordinates multiple LLMs dynamically, including recursive instances of itself, and decides on each request whether to answer directly or assemble a team. The release headline is “One Model to Command Them All.”

Fable 5 and Mythos are different. They are single Anthropic models. Fable 5 is Anthropic’s most powerful generally available model, a “Mythos-class” model made safe for public use, a tier above Opus 4.8. Mythos Preview, released April 7, 2026, is the frontier model Anthropic described as too dangerous to release. One important naming detail: Sakana named the older Mythos Preview in its comparison, not the current Mythos 5. We cover the Anthropic side in Fable 5 vs Mythos 5 and the Mythos-class model explained.

So the matchup is a model-of-models against single models. Keep that frame. It changes how you read every number below.

Fugu and Fugu Ultra, briefly

Fugu ships in two variants through one endpoint. “Fugu” is the balanced, low-latency option for everyday work, coding, code review, chatbots, and interactive services. “Fugu Ultra” targets maximum answer quality for AI research, paper reproduction, cybersecurity analysis, and literature or patent investigation. The beta and much of the press called the small variant “Fugu Mini,” but the release page leads with “Fugu” and “Fugu Ultra,” so those are the names to use.

The honest spine: orchestrator vs single model

Here is the part you cannot skip. Fugu is an orchestrator. When it produces a strong answer, it may have done so by calling another vendor’s frontier model and synthesizing the result. That includes calling Opus 4.8, calling Gemini, or calling recursive copies of itself.

So when you see a result where Fugu “beats Opus 4.8,” it is entirely possible that Fugu reached that result by calling Opus 4.8 and adding a verification or synthesis layer on top. That is a real capability, and it can be genuinely useful. It is not the same thing as a single model beating Opus on its own weights.

Fable 5 and Mythos are single models. They answer from their own parameters. There is no team behind the curtain.

This is why we will never write “Fugu beats Fable 5” in this article, and you should be skeptical of any source that does. The categories are different. A fair reading is “an orchestrated system reached frontier-comparable quality, partly by routing to frontier models.” That sentence is less punchy and far more accurate. If you want the full benchmark walkthrough, see Sakana Fugu benchmarks.

Tier one: parity with Fable 5 and Mythos Preview

Sakana’s first claim is parity. Per Sakana, Fugu Ultra stands shoulder-to-shoulder with Fable 5 and Mythos Preview across engineering, scientific, and reasoning benchmarks. Read that carefully. “Shoulder-to-shoulder” is a parity word. It does not say Fugu Ultra wins. It says Fugu Ultra keeps pace.

Two things make this claim more interesting than a raw score.

First, the peer Sakana chose. Mythos Preview is the April frontier model, not the current Mythos 5. Pricing reflects that gap: Anthropic lists Fable 5 and Mythos 5 at $10 per million input tokens and $50 per million output, while Mythos Preview sat at $25 input and $125 output (Anthropic pricing, June 9, 2026). Naming the older preview is a defensible choice for a reproducible comparison, but it is worth stating plainly so you are not comparing against a model that no longer represents the current ceiling.

Second, the mechanism. If Fugu Ultra reaches Fable-5-level reasoning by orchestrating a team, the parity is real at the system level even if no single underlying model matched Fable 5 alone. That is the honest framing. For the single-model side of this matchup, Claude Fable 5 vs Opus 4.8 walks through where Fable 5 itself lands.

Tier two: where Sakana claims Fugu outperforms

This is a separate claim, against a separate set of models, and it deserves its own section so it never blurs into the parity claim above.

Per Sakana, Fugu “consistently outperforms” three frontier models on specific applications:

The applications named are narrow and concrete: AutoResearch, Rubik’s Cube, Mechanical Design, Japanese Handwriting Analysis, One-Shot Chess, and Financial Time Series Prediction.

Two honest notes apply here too. These are application-level wins, not general-benchmark wins. An orchestrator can score high on a structured, multi-step task like AutoResearch or chess because it can plan, delegate, verify, and retry. That is exactly where a coordination layer earns its keep. And again, the outperformance may come partly from routing to the same frontier models it is being compared against. A “beats Opus 4.8 on Rubik’s Cube” result that was reached by calling Opus inside the loop is a system win, not a weights win.

So the accurate summary is: Fugu’s coordination layer adds measurable value on structured, verifiable tasks, sometimes enough to top a single frontier model on that specific task.

The comparison table

This table keeps the category row first on purpose. Read that row before the rest.

Dimension Fugu / Fugu Ultra Fable 5 Mythos (Preview / 5)
What kind of system Orchestrator: trained conductor that calls multiple LLMs, including itself Single Anthropic model Single Anthropic model
Vendor Sakana AI Anthropic Anthropic
Sakana’s claim vs this model Parity (“shoulder-to-shoulder”) with Fable 5 and Mythos Preview Named parity peer Named parity peer (Preview, not 5)
Separate outperform claim Vs Gemini 3.1 Pro, Opus 4.8, GPT 5.5 on named apps Not the outperform target Not the outperform target
Pricing (sourced) Reported tiers + PAYG, all [VERIFY] $10 in / $50 out per 1M Preview $25 in / $125 out; Mythos 5 $10 / $50
API surface One OpenAI-compatible endpoint, both variants Anthropic API Anthropic API
Strength Structured multi-step tasks, governance routing General-purpose frontier quality, GA-safe Raw frontier ceiling

The pricing numbers for Fugu are reported and not from the release page itself, so treat every Fugu dollar figure as unverified until you confirm it live. The Anthropic figures are sourced from Anthropic’s June 9, 2026 pricing. For a closer look at Fable 5’s own scores, see Claude Fable 5 benchmarks.

Pricing, gated honestly

Sakana confirmed a pricing structure on the release page: subscription tiers for everyday use, plus a pay-as-you-go plan for heavier and enterprise workloads. That concept is solid.

Every actual number is not. As of 2026-06-22, the reported figures come from JS-rendered or secondary sources, not the release page. Reported subscription tiers run $20, $100, and $200 per month for both models, with a launch promo of a free second month if you subscribe before the end of July 2026. Reported pay-as-you-go pricing is roughly $5 input, $30 output, and $0.50 cached per million tokens, with a context surcharge above 272K tokens. The base “Fugu” variant is reportedly passthrough-billed at the standard rate of the underlying model it calls. No standalone free tier surfaced.

Reported, verify live as of 2026-06-22. Do not plan a budget on these numbers until you confirm them in your console. The structure is real; the digits are not yet.

The research lineage, and what it does not prove

Sakana did not invent orchestration. Mixture-of-Agents from Together AI (ICLR 2025) already showed orchestrated models beating a single model. Fugu’s narrower novelty is a learned, adaptive, cost-selective topology shipped as one endpoint, backed by a trained-conductor research line.

That line is real and citable. Two ICLR 2026 papers sit behind the approach. Trinity, “An Evolved LLM Coordinator” (arXiv:2512.04695), is a sub-20K-parameter coordinator optimized by derivative-free evolution, with Thinker, Worker, and Verifier roles. Conductor, “Learning to Orchestrate Agents in Natural Language” (arXiv:2512.04388), is a 7B model trained with reinforcement learning that learns the communication structure and claims to beat Mixture-of-Agents at lower cost.

These are different methods and different sizes. Evolution versus RL, sub-20K versus 7B. Do not conflate them, and do not assume either set of specifics maps cleanly onto the shipped product. The official release gives no product parameter count, so applying the 7B details to Fugu itself is third-party inference, not a stated fact.

What separates Fugu from neighbors is concrete. Routers like OpenRouter or Martian pick one model and send your request there. Agent frameworks like Swarm, AutoGen, or LangGraph make you the coordinator. Fugu trains the coordinator and hides it behind a single call.

How this fits your Apidog workflow

Fugu exposes one OpenAI-compatible endpoint. You point an existing OpenAI client at it with your key, and you do not migrate an SDK. Both variants run through the same endpoint, and Fugu decides whether to answer directly or assemble a team.

The base URL is not published on any public page as of 2026-06-22, so do not trust any host you see floating around. Copy the real base URL from your console at console.sakana.ai, then drop it into a standard OpenAI request. The model id strings reported are fugu and fugu-ultra, with a possible dated form; confirm the exact id in your console rather than hardcoding a dated string.

Here is the shape of a call, using a placeholder you replace with your real base URL:

from openai import OpenAI

client = OpenAI(
    api_key="YOUR_SAKANA_API_KEY",
    base_url="<YOUR_FUGU_BASE_URL_FROM_CONSOLE>",  # copy from console.sakana.ai
)

response = client.chat.completions.create(
    model="fugu-ultra",  # confirm exact id in console; "fugu" for the balanced variant
    messages=[
        {"role": "system", "content": "You are a careful code reviewer."},
        {"role": "user", "content": "Review this pull request for security issues."},
    ],
)

print(response.choices[0].message.content)

Because it speaks the OpenAI chat completions format (OpenAI API reference), you can test Fugu the same way you test any other model endpoint. In Apidog, create a request against your console base URL, set the model field to fugu-ultra, and save it as a reusable case. You can run the same request against Fable 5 or Opus 4.8 endpoints side by side, compare responses, and assert on the parts that matter to you. That is the practical way to check Sakana’s parity claim with your own prompts instead of trusting a marketing table. Download Apidog to set up the comparison.

One operational detail worth a callout for compliance teams. Fugu’s agents are swappable, you can opt specific agents out of the pool for data or compliance reasons, and Sakana says Fugu dynamically routes around provider restrictions. If you are testing in a regulated context, exercise that opt-out path and assert that excluded providers never appear in the response trace.

The verdict, both sides steelmanned

The case for being impressed: Sakana shipped a trained conductor as one clean endpoint, backed by a real research line, and claimed both frontier parity and application-level wins. On structured, verifiable tasks like AutoResearch and chess, a coordination layer is exactly the right tool, and early users back this up. Sakana reports a software engineer who said Fugu Ultra surfaced more than twenty issues in a code review where other tools flag about three, calling it better than GPT-5.5. A security engineer said one scoped instruction drove a full end-to-end assessment, including recon and an auth review, while staying in scope.

The case for restraint: the parity claim is against the older Mythos Preview, not Mythos 5, and the outperform claim sits in a separate tier against a different model set. Both claims come from a system that can reach its numbers by calling the same models it is compared to. Mixture-of-Agents already proved orchestration can beat single models, so the high-level idea is not new. The honest read is “frontier-comparable orchestration with a learned conductor and a strong governance story,” not “the new best model.”

For most teams, the right move is neither hype nor dismissal. Run Fugu Ultra against Fable 5 and Opus 4.8 on your own structured tasks, watch where the coordination layer earns its cost, and verify every pricing number before you commit. Sakana’s branding leans into the metaphor: fugu is the pufferfish, safe only when a skilled chef prepares it, and orchestration is that careful preparation. Whether the preparation is worth the markup is a question your own test suite can answer faster than any release page.

button

Frequently Asked Questions

Does Fugu Ultra beat Fable 5?

No, and Sakana does not claim it does. Per Sakana, Fugu Ultra stands shoulder-to-shoulder with Fable 5 and Mythos Preview, which is a parity claim, not a win. Because Fugu is an orchestrator that can call frontier models, any “win” you see may come from routing to those models. See Fable 5 vs Mythos 5 for the single-model side.

What does Sakana mean when it says Fugu outperforms Opus 4.8?

That is a separate claim from the parity claim, and it applies to specific applications, not general benchmarks. Per Sakana, Fugu consistently outperforms Gemini 3.1 Pro, Opus 4.8, and GPT 5.5 on tasks like AutoResearch, one-shot chess, and financial time-series prediction. Keep in mind Fugu may reach those results by calling Opus inside its own loop, so it is a system win, not a single-model win.

Why does Sakana compare against Mythos Preview instead of Mythos 5?

Mythos Preview is the April 2026 frontier model Anthropic called too dangerous to release, while Mythos 5 is the current generally available version. Sakana named the older preview in its comparison. It is a defensible choice for a reproducible test, but it means the parity claim is not measured against today’s ceiling. The Mythos-class model explained covers the distinction.

Is Fugu a single model or a group of models?

It is a group. Fugu is a trained conductor that delegates to multiple LLMs, including recursive copies of itself, and presents the whole system as one model behind one OpenAI-compatible API. Fable 5 and Mythos are single Anthropic models that answer from their own weights, which is the core reason this comparison crosses two different system categories.

How do I test Fugu against Fable 5 myself?

Point an OpenAI-compatible client at your Sakana console base URL, set the model to fugu-ultra, and run the same prompts you run against Fable 5 or Opus 4.8. In Apidog you can save each model as a request, run them side by side, and assert on the output that matters to you. That turns Sakana’s parity claim into something you measure instead of trust.

How much does Fugu cost compared to Fable 5?

The pricing structure is confirmed (subscription tiers plus pay-as-you-go), but every Fugu dollar figure is reported from secondary sources and unverified as of 2026-06-22, so confirm it in your console before budgeting. For reference, Anthropic lists Fable 5 at $10 per million input tokens and $50 per million output. Our Sakana Fugu benchmarks piece tracks pricing as it gets confirmed.

button

Explore more

What's New in Gemini 3.7 Flash? Features, Benchmarks, and API Access

What's New in Gemini 3.7 Flash? Features, Benchmarks, and API Access

What's new in Gemini 3.7 Flash? Full benchmark table vs 3.6, intro API pricing, coding and agent gains, every access channel, and a 5-minute cURL test.

14 August 2026

ChatCompletions vs Anthropic Messages vs Responses API: Testing DeepSeek V4 Pro's Three API Formats

ChatCompletions vs Anthropic Messages vs Responses API: Testing DeepSeek V4 Pro's Three API Formats

DeepSeek V4 Pro speaks three API formats: OpenAI ChatCompletions, Anthropic Messages, and its own Responses API. Compare request shapes with real examples and test all three side by side in Apidog.

13 August 2026

DeepSeek API Price Increase Is Coming: A Developer's Cost-Optimization Playbook

DeepSeek API Price Increase Is Coming: A Developer's Cost-Optimization Playbook

DeepSeek says a significant API price increase is coming. Cut your exposure now: prompt caching, Flash/Pro routing, off-peak batching, failover testing, and what 1.5x-3x scenarios do to your bill.

13 August 2026

Practice API Design-first in Apidog

Discover an easier way to build and use APIs

Fugu Ultra vs Fable 5 vs Mythos: How the Sakana Orchestrator Stacks Up Against the Frontier