One Model Family, Two Gold-Level Results: Fine-Tuning Nemotron for IOI and IMO
NVIDIA's Nemotron 3 foundation model has been fine-tuned to achieve gold-medal level results in both the International Olympiad in Informatics (IOI) and the International Mathematical Olympiad (IMO) 2026.
Intelligence analysis by Gemini 2.5 Flash

Researchers successfully specialized Nemotron 3, a large language model, for two distinct and highly challenging competitions: IOI, which demands algorithmic coding, and IMO, requiring rigorous mathematical proofs. This was achieved through a reusable recipe involving supervised fine-tuning (SFT), reinforcement learning (RL), and feedback-driven inference, demonstrating the model's ad…
Imagine a super-smart robot brain, called Nemotron, that's really good at learning. Scientists taught this brain how to become a champion in two very different school contests: one for writing computer code to solve puzzles, and another for solving really hard math problems and explaining the answers like a teacher. It's like teaching a brilliant student to not just be good at one subject, but to become a top expert in two completely different, tough challenges by practicing a lot and learning from its mistakes.
Analysis
The recent achievements with Nemotron 3 underscore a significant advancement in the field of artificial intelligence, particularly in the specialization of large language models. By demonstrating gold-medal level performance in both the International Olympiad in Informatics (IOI) and the International Mathematical Olympiad (IMO) 2026, NVIDIA's research team has showcased a powerful, adaptable framework for creating highly capable specialist AI systems. This success was not merely a result of a powerful base model but a testament to a carefully designed, reusable specialization recipe that combines advanced training methodologies with sophisticated inference strategies.
Nemotron 3
Nemotron 3 serves as the robust foundation for these specialized models, proving its versatility across vastly different intellectual challenges. The core idea was to take a strong general-purpose model and adapt it to demanding domains rather than building new foundation models from scratch for each task. This approach involved starting with Nemotron 3, curating extensive domain-specific datasets, and applying standard post-training methods like Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL). The results indicate that Nemotron 3, even at different scales (e.g., Nemotron-3-Nano-CC and Nemotron-3-Ultra-CC), can be effectively molded to achieve peak performance in highly specialized areas, validating the concept of a 'reusable specialization recipe.'
IOI 2026
For the International Olympiad in Informatics, Nemotron-3-Ultra-CC achieved an impressive score of 535.4 out of 600, significantly surpassing the gold threshold of 361.12 and even the top human score of 498.27. This result was obtained under live, prospective run conditions, mirroring the constraints faced by human contestants. The specialization process for IOI involved curating 22,000 competitive programming problems and generating synthetic reasoning traces to train two specialists. The iterative generate-evaluate-refine strategy, dubbed GenCorrect, played a crucial role, turning the gains from fine-tuning into larger improvements over multiple feedback rounds. This demonstrates that combining a capable specialist model with an effective inference loop is key to pushing performance boundaries in complex coding challenges.
IMO 2026
In the International Mathematical Olympiad, the Nemotron system scored 30 out of 42 points, exceeding the official gold threshold of 29. This included full credit on four of the six problems, a remarkable feat given the competition's demand for rigorous natural-language proofs. The IMO project applied a similar specialization idea, training specialists with SFT and RL on a corpus of nearly 415,000 quality-filtered examples across 15,818 unique proof problems. The training data taught the model not just final answers but also proof generation, refinement, verification, and meta-verification. The final system leveraged the complementary strengths of both SFT and RL checkpoints, generating candidate proofs, scoring them, producing critiques, and refining promising attempts, all within a natural language framework without external tools or internet access.
Key points
- NVIDIA's Nemotron 3 achieved gold-medal level results in both IOI 2026 and IMO 2026.
- The success was attributed to a reusable specialization recipe involving SFT, RL, and feedback-driven inference.
- For IOI, Nemotron-3-Ultra-CC scored 535.4/600, surpassing the top human score and gold threshold.
- For IMO, the system scored 30/42, exceeding the official gold-medal threshold with rigorous natural-language proofs.
- The approach emphasizes co-designing the model, data, and inference loop for optimal performance in demanding domains.
This breakthrough suggests that AI models can be reliably trained to tackle highly complex, creative problem-solving tasks, potentially accelerating scientific discovery, improving educational tools, and enabling AI assistants capable of advanced reasoning in various professional fields.
The extensive computational resources and specialized data curation required for such gold-level performance might limit the accessibility and widespread application of these highly specialized AI systems, potentially widening the gap between well-resourced and less-resourced AI development teams.



