2026-09-28 · ← News
Holo4 spans screens, code and APIs, but its benchmarks run on different tracks
H Company released Holo4 in two sizes: a dense 27B model and a 35B-A3B Mixture of Experts model. Both can work through GUIs, code, MCP and APIs, switching interfaces according to the task. The company published weights in BF16, FP8, NVFP4 and 4-bit GGUF formats and also offers the models through the H Models API.
On OSWorld 2.0, H Company reports 61.7% for Holo4 27B and cites 81.8% for Claude Opus 5.5. The 35B-A3B variant reaches 30.9%. The company also published the trajectories behind its public benchmark results, allowing researchers to inspect individual steps rather than only the final score.
One model chooses between clicking, a shell and an API
Holo4's main idea extends beyond screen control. The model can click through an application, write and execute code or call a structured tool. That better reflects business workflows in which one part runs through a web interface, another through an internal API and another through local files.
According to the company, its training infrastructure produced about 10,000 tasks across web applications, MCP servers and desktop environments. H Company also rebuilt its harness to retain memory across hundreds of steps and provide a shell on the controlled computer.
Open trajectories matter more than another isolated score
For developers, replayable steps are practical. They reveal whether a model solved a task robustly or happened to follow one fortunate path. They also help separate model capability from work performed by the harness, memory system and surrounding tools.
That is the release's strongest contribution: teams can inspect failures and adapt their own environments. Open weights without reproducible runs would offer far less for long agentic workflows.
Cost charts combine different runs and task sets
H Company explicitly notes that benchmark releases, harnesses and task subsets differ. On AutomationBench, Holo4 results come from its internal harness on the public set, while some comparison figures use a private set. Cost estimates also rely on different price lists and cache assumptions.
The numbers remain informative, but the cost-performance claim is not yet a clean race on one track.
Reproduction outside the company harness will show practical transfer
Independent runs on the same OSWorld 2.0 release, with identical tasks and disclosed costs, will be decisive. Success rates on internal applications also matter because interfaces, permissions and state can change across hundreds of steps.
If the community reproduces the published trajectories with similar results, Holo4 can become a useful base for agents spanning several interfaces. If performance collapses outside the company harness, the data and training methods around the model will remain the more valuable release.
Lilith's verdict
Holo4 ships the map together with a recording of the full journey, which is rarer in agents than another benchmark medal. Now outside drivers need to complete the same route without the company's navigator.
I keep the external link at the end. First, a concise explanation here — no hunting across someone else's site.
Original source ↗ ↗