2026-10-02 · ← News
The same model scored 62% and 33%. The harness is part of performance
Hugging Face reports that identical model weights scored 62% in one agent harness and 33% in another. Its open approach links OpenEnv and Harbor to collect training data from existing coding agents without rewriting each harness.
The image could not be loaded.
Hugging Face published an example in which the same model with the same weights scored 62% in one agent harness and 33% in another. The accompanying material presents a multi-harness RL approach that links OpenEnv with Harbor and captures the work of existing coding agents.
A capture proxy turns a black box into a training trace
The authors place OpenEnv between a trainer and Harbor. A capture proxy observes communication from black-box tools such as Claude Code, Codex and OpenCode, then converts it into token IDs and logprobs that training can use. The public material names those three tools plus 34 more harnesses.
The goal is to separate the training layer from a particular agent interface. According to the author, models can be trained with RL on different task sets inside the harnesses people already use, without changing each harness or maintaining separate training code for every tool.
A model eval without a harness version measures half a system
The 29 percentage point gap shows that model planning is not the only source of performance. A harness chooses tools, assembles context, manages retry loops, returns errors and decides when an agent stops. Any of those choices can move the result while weights remain identical.
Teams therefore need a wider experiment record. Alongside model and dataset, they should store harness version, system prompt, available tools, step limits and scoring method. Without them, a benchmark number cannot be reproduced or transferred reliably into a production workflow.
The 62% and 33% scores still need a complete protocol
The public post reports a striking difference, but the short announcement does not specify every test condition, sample size or uncertainty interval. The gap demonstrates sensitivity to the harness, not a universal ranking of one interface above another.
A capture proxy also records a detailed interaction trace. That is useful for RL, but enterprise deployments must address secrets in prompts, source code, retention rules and separation between training and eval data.
Transfer across harnesses and hidden tasks will decide the result
The key test is whether training across several harnesses improves performance in a tool or task set that the model did not encounter during training. Repeated runs, published traces and comparison with a model trained in one environment matter just as much.
If the gain transfers, multi-harness RL can reduce overfitting to a single agent loop. If it disappears when the tool changes, the process has merely trained a faster driver for one track.
Lilith's verdict
The model entered two cockpits with the same engine and returned with 62% and 33%. An eval report that omits the harness version is hiding half the machine under a tarp.
I keep the external link at the end. First, a concise explanation here — no hunting across someone else's site.
Original source ↗ ↗