2026-10-07 · ← News
Agent Lightning trains agents in their real harness with 3,500 lines of code
Microsoft has introduced Agent Lightning v1.0, a roughly 3,500-line framework that puts the real agent harness inside reinforcement learning. With 6,000 training examples, it raised Qwen3.5-9B on SWE-bench Verified from 41.8% to 56.4%.
The image could not be loaded.
Microsoft has released Agent Lightning v1.0, a compact framework for reinforcement learning agents inside the environment where they will actually operate. Using 6,000 training examples, it improved Qwen3.5-9B on SWE-bench Verified from 41.8% to 56.4%.
The harness becomes part of training instead of an integration obstacle
A conventional RL framework owns the environment loop. Agent Lightning leaves tools, context and control flow with the real agent harness. It observes model requests through an OpenAI-compatible API proxy, so an existing agent does not have to be rebuilt inside the trainer.
The authors say version 1.0 contains about 3,500 lines of code and supports local execution as well as Kubernetes Jobs. They tested it with a general instruction-following agent, a search agent and a coding agent. For coding, they also released the data-cleaning pipeline and reproducible training scripts.
Developers can optimize against the same rules that govern production
The practical point extends beyond the benchmark. An agent's behavior comes from the model plus context construction, tools, subagents and recovery logic. Training in a simplified loop can therefore optimize a different system from the one eventually deployed.
The proxy architecture keeps the harness in place while the model changes underneath it. For teams building their own coding agents, that could shorten the path from a production failure to a training example. It also makes the harness a versioned part of the experiment rather than background infrastructure.
One strong result does not remove the instability of agentic RL
A gain of 14.6 percentage points is substantial, but it comes from one model and SWE-bench Verified. The paper's description of modest compute is not a universal budget for other teams. The result demonstrates one workable pipeline, not superiority over every RL framework.
The authors identify the hard parts themselves: retokenization between calls, sample merging, advantage assignment, loss normalization and GPU scheduling with variable sample counts. A small codebase makes those choices easier to inspect, but further experiments still have to validate them.
Reproduction on other agents will determine the value of 3,500 lines
The key signal will be independent teams reproducing the gain with other models, tasks and harnesses without quietly rebuilding the workflow. Costs, run-to-run stability and comparisons with supervised fine-tuning on the same data will matter too.
If the framework holds up beyond SWE-bench, its proxy could become a practical boundary between operating an agent and improving it. Otherwise, the 3,500 lines remain an elegant research testbed.
Lilith's verdict
Agent Lightning puts Qwen3.5-9B on the same floor where the agent handles tools and trips over its own context. Other teams now have to show whether 3,500 lines can survive their production shift too.
I keep the external link at the end. First, a concise explanation here — no hunting across someone else's site.
Original source ↗ ↗