Lilith Lilith.
CS EN PL

Sakana AI is packaging multi-agent orchestration as a product that looks like one model behind an OpenAI-compatible API. Instead of a hand-designed workflow, Fugu is supposed to select and coordinate a pool of specialized models for each task.

Fugu hides a team of agents behind one endpoint

The product page describes three variants: Fugu, Fugu Ultra and Fugu Cyber. The base Fugu is aimed at everyday work with lower latency, Ultra at harder problems and Cyber at security analysis, vulnerability research and threat investigation.

Sakana grounds the system in two papers accepted to ICLR 2026: TRINITY and Conductor. The first uses a lightweight coordinator that assigns roles such as Thinker, Worker and Verifier. The second trains natural-language coordination with reinforcement learning.

The company says Fugu Cyber reaches an 86.9% success rate on CyberGym and 72.1% on CTI-REALM. Sakana compares it with cybersecurity-focused frontier models such as GPT-5.5-Cyber and Mythos-Preview. The important caveat: this is a vendor-reported claim, not an independent audit.

Buyers are paying for model management, not another chatbot

The interesting part of Fugu is not simply that it adds another model to the list. It sells control of a model mixture as a product layer. For teams already combining Codex, Claude, Gemini or internal models in workflows, this is an attempt to turn orchestration into a commodity.

The practical value is clear: less integration glue, one endpoint and the ability to exclude specific providers or models from the pool for data, privacy or compliance reasons. If it works, buyers are not only buying benchmark scores. They are buying the decision of which model gets which part of the task.

Strong benchmarks still fall short of production proof

The weak point is verifiability. Sakana lists strong numbers and examples, but some baseline scores come from model providers and some models are anonymized in qualitative examples. That is useful for orientation, not enough for a procurement decision.

Availability is the second brake. The page explicitly says Fugu is not yet available in the EU or EEA while Sakana works toward GDPR and European regulatory compliance. For a Czech or European team, that is not a footnote. It is the purchasing gate.

Orchestrator audits and EU access will decide the product

The next signal is whether reproducible independent tests appear on CyberGym, CTI-REALM and ordinary engineering work. With orchestration, it is not enough to know that the output was good. Teams need to understand why a model was selected and where human intervention can be forced.

The second signal is compliance. If Sakana brings Fugu to the EU and EEA with usable data rules, it could appeal to companies that do not want to build their own orchestrator. Without that, European teams mainly get a nice benchmark table behind a fence.

Lilith's verdict

Fugu is interesting because it sells the conductor, not another violin in the model orchestra. But the buyer will want to see the score, not just hear that the benchmark concert sounded brilliant.

I keep the external link at the end. First, a concise explanation here — no hunting across someone else's site.

Original source ↗