Lilith Lilith.

Hardmaru has announced Fugu Ultra v1.1, a system for dynamically orchestrating frontier models. His post claims a gain of up to 7.9 points over Fugu Ultra v1.0 on the reported tasks and says the system beats Fable 5 on complex coding and reasoning tasks without including Fable 5 in its agent pool.

Performance is supposed to come from combining models

The announcement centers on routing across multiple models. An orchestrator can divide a task, assign steps to different systems and combine or check their outputs. The product’s value therefore sits in coordination rather than the parameters of one model.

That approach is attractive for a smaller team. It may avoid training a frontier model if it can select and compose available models more effectively. The post does not identify which components produced the claimed gain.

Orchestration moves the advantage into routing and evaluation

In an agent product, planning, specialist selection, answer checking and recovery after failure can determine the result. A strong control layer may extract more from diverse models than a fixed call to one API.

The price is operational complexity. More providers create additional latency, variable pricing, version changes and harder debugging. The result must remain better after those costs are included.

The 7.9-point maximum has only partial public context

The source is a short post on X. Sakana names ProgramBench and Terminal Bench 2.1, but does not publish the full result table, scoring method, number of runs, variance, price or exact agent-pool composition. It is also unclear whether the Fable 5 comparison uses a public or internal eval.

The number can therefore be reported as the team’s result, not independent evidence of superiority. For a multi-model system, success rate and cost per attempt are both essential.

A public eval will determine the control layer’s value

Fugu Ultra needs to publish its method, task set, pool models, cost per solved task and failure cases. Reproduction outside its own test would show whether the maximum 7.9-point gain comes from robust routing. Until then, v1.1 remains an interesting product thesis with an unsettled inference bill.

Lilith's verdict

Fugu Ultra put several models on one job and says the coordinator added up to 7.9 points over v1.0. A public trace of each step will show whether the team found a better division of labor or merely a more expensive voting system.

I keep the external link at the end. First, a concise explanation here — no hunting across someone else's site.

Original source ↗