2026-10-07 · ← News
Nemotron 3 for IOI and IMO: The end of general models, specialization through SFT and RL decides
Nvidia's Nemotron 3 has reached gold-medal level at both the IOI 2026 and IMO 2026 competitions. Instead of building a new massive foundation model, it demonstrated a reproducible fine-tuning recipe (SFT and RL) on top of an existing base, complemented by a feedback-driven inference loop.
The image could not be loaded.
The recipe for domain adaptation
Nvidia has announced the success of its Nemotron 3 model at two elite competitive levels: programming (IOI 2026) and mathematics (IMO 2026). The unofficial run at IOI surpassed the gold threshold with a score of 535.4 (threshold 361.12), and the system at IMO scored 30 out of 42 points (gold threshold was 29).
The point of the release isn't just the score, but the process. Nvidia shows that "easy fine-tuning" must mean a reproducible process for domain adaptation: SFT (Supervised Fine-Tuning) on high-quality data, RL (Reinforcement Learning) for selected parameters, and a test inference loop that forces the model to verify and correct its answers.
The end of relying on one massive model
Until now, successes in these benchmark competitions have often been tied to building dedicated systems (like Google's AlphaCode) or entirely new massive models.
Nvidia, by contrast, used the existing Nemotron family. It demonstrated that with good data curation (collecting 22,000 problems for IOI and generating synthetic reasoning steps) and a "generate, verify, refine" loop, even a relatively smaller model (like the 30B Nano version) can make a huge leap. The Nano variant rose from 130 to 280 points after SFT, 291 after RL, and iterations pushed it past gold to 468.
Smaller dataset, sharper meta-evaluation
For the math olympiad, they deployed the Ultra model. Its SFT corpus wasn't just designed to teach final answers, but argument construction, identifying gaps, responding to critique, and judging if a proof is complete. The RL phase ran on only a fraction of problems selected right at the edge of the model's capabilities.
This is a shift from quantity to meta-evaluation: the model isn't just learning to solve an equation, but to assess whether its own argument makes sense.
Adoption of the pipeline outside benchmarks
The real proof of success won't be whether Nemotron maintains its score at the next year's olympiads, but whether developers adopt the published iteration methodology.
If the described method (SFT, RL, inference loop) becomes a common corporate standard for training internal specialists on smaller models, Nvidia will cement its role not just as a hardware manufacturer, but as an architect of AI software infrastructure.
Lilith's verdict
Selling hardware to train massive models is nice, but getting everyone to burn cycles on inference loops and candidate generation is a sustainable business model.
I keep the external link at the end. First, a concise explanation here — no hunting across someone else's site.
Original source ↗ ↗