AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: What It Takes To Fine-Tune Nemotron For IOI And IMO on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get tech for your team delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

Hugging Face says two specialized systems built from its Nemotron 3 models scored 535.4 out of 600 at the 2026 IOI and 30 out of 42 at the 2026 IMO. The IMO proofs were officially graded; the IOI score came from an unofficial run and did not count toward the competition ranking.

Hugging Face says two specialized systems built from its Nemotron 3 model family reached gold-medal-level scores at the 2026 International Olympiad in Informatics (IOI) and International Mathematical Olympiad (IMO), as detailed in the original analysis. The company reported 535.4 points out of 600 at the IOI and 30 out of 42 at the IMO, but the results have different status: IMO graders officially evaluated the proofs, while the IOI score came from an unofficial run that was excluded from the contest ranking.

For the IOI, Hugging Face used a competition-specific version of Nemotron-3-Ultra-CC, trained with supervised fine-tuning (SFT). The system also used GenCorrect, an iterative process that generates candidate code, evaluates it and revises solutions. Hugging Face says the run was conducted prospectively under the competition’s time, internet-access and submission constraints. Its reported 535.4-point score was above the stated gold threshold of 361.12 and the top human score of 498.27, but it was not an official competition entry or ranking result.

For the IMO, the team combined the general Nemotron 3 Ultra model with SFT and reinforcement-learning checkpoints in a natural-language proof system. It generated candidate proofs, scored and critiqued them, then revised selected attempts. Hugging Face says official graders awarded the submitted proofs 30 of 42 points, above the stated gold threshold of 29, with full marks on four of the six problems. The company says the system used no formal prover, external tools or internet access.

The projects used separate specialist training data. The IOI work drew on 22,000 programming problems and synthetic reasoning traces. The IMO SFT dataset contained 414,890 quality-filtered examples from 15,818 proof problems; its reinforcement-learning model was trained on 9,597 problems selected near the model’s capability frontier. These figures and results are reported by Hugging Face.

At a glance
reportWhen: Reported after the 2026 IOI and IMO res…
The developmentHugging Face reported gold-level results for Nemotron-based systems at the 2026 IOI and IMO, using fine-tuning and iterative solution-refinement methods.
At a glance
reportWhen: Reported after the 2026 competitions
The developmentHugging Face reported that systems fine-tuned from Nemotron 3 scored above the gold thresholds at IOI 2026 and IMO 2026.

Why the Two Scores Differ

The report offers evidence that a shared model family can be adapted for distinct technical contests through domain-specific fine-tuning and additional computation at answer time. The systems did not rely on one method alone: the IOI setup used code generation and iterative evaluation, while the IMO setup combined different model checkpoints with proof critique and revision. That distinction matters because programming submissions must pass tests, including hidden tests, while mathematics submissions must present arguments that graders judge as valid.

The scores should not be treated as equivalent evidence. The IMO result received official grading, giving it a direct connection to the contest’s scoring process. The IOI score is a reported competition-like benchmark run, not a medal or official placement. Both results concern narrowly defined olympiad tasks; they do not by themselves show how the models perform across everyday software development, general mathematics or unfamiliar evaluation settings.

For researchers, the work also highlights a practical question: how much performance comes from the underlying model, the specialist data, or repeated generation and checking at inference time? Hugging Face’s report describes a combination of these elements, rather than isolating the contribution of each one. Reproducible datasets and benchmarks could help others test that account.

From IOI Experiments to Proofs

Hugging Face presents the 2026 work as a continuation of its experiments at the 2025 IOI, where it tested whether post-training and additional solution search could improve open-weight model performance. The company reported that a Nemotron-3-Nano-CC model rose from 130 points before post-training to 280 after SFT and 291 after reinforcement learning. With GenCorrect, it reached 468, above the stated 2025 gold threshold of 438.3. An Ultra-CC version scored 502 using the same test-time strategy.

The newer projects apply related ideas to two different tasks. The IOI system writes programs that need to meet contest constraints and pass tests; the IMO system produces written mathematical proofs. According to Hugging Face, the IMO work found useful differences between its SFT and reinforcement-learning checkpoints, leading the team to combine them with the general model rather than depend on a single specialist checkpoint.

The company says it is releasing an IMO collection with SFT and reinforcement-learning checkpoints, both training datasets, and Nemotron-IMO-Bench, a benchmark of 200 olympiad-level problems. It also refers to an IMO paper and a NeMo-Skills repository. These materials may help researchers examine the proof system and evaluate it on further problems, though release details in the source account are incomplete.

“Success at both points to something broader.”

— Hugging Face

Validation and Generalization Gaps

The central limitation is the unofficial status of the IOI run. Hugging Face describes it as prospective and conducted under contest-like constraints, but the supplied account does not explain how the run was audited or provide independent verification. The score was not included in the official ranking, so it should be described as a company-reported benchmark result, not an official medal.

The IMO proofs were officially graded, but the report does not establish how performance would hold across other proof styles, unseen problems or broader mathematical work. Nor does it isolate the effect of training data, fine-tuning, checkpoint combination and iterative inference. The source also does not provide an independent replication of either system’s results or a timetable for all cited materials.

Releases and Independent Tests

Hugging Face says its IMO collection includes model checkpoints, training datasets and the 200-problem Nemotron-IMO-Bench, alongside a paper describing the training and generate-verify-refine approach. Researchers will be able to use those materials to inspect the methods and test performance beyond the reported contest submissions. The company has not provided a timetable for further releases in the supplied account.

The most useful next steps would be external evaluations of the IOI procedure, independent replication of the scores and tests on additional tasks. Until then, the officially graded IMO score and the unofficial IOI benchmark score remain separate findings, with distinct limits on what each confirms.

Key Questions

What scores did the Nemotron systems report?

Hugging Face reported 535.4 out of 600 at the 2026 IOI and 30 out of 42 at the 2026 IMO. The IOI figure came from an unofficial run; IMO graders officially evaluated the proofs.

Did Nemotron win an official IOI medal?

No. The company says the IOI system’s run was unofficial and did not count toward the official ranking. Its score exceeded the stated gold threshold, but that does not make it an official medal result.

How did the IMO system produce its proofs?

It combined the general Nemotron 3 Ultra model with supervised fine-tuning and reinforcement-learning checkpoints. The system generated candidate proofs, evaluated and critiqued them, then revised selected attempts, according to Hugging Face.

What materials does Hugging Face say it is releasing?

The company says its IMO collection includes SFT and reinforcement-learning checkpoints, both training datasets and Nemotron-IMO-Bench, a benchmark of 200 olympiad-level problems. It also points to an IMO paper and a NeMo-Skills repository.

Do the scores show broad performance across math and programming?

Not on their own. The results cover specific olympiad tasks, and the IOI score has no official ranking status. Independent testing would be needed to establish how well the systems generalize to other contests or real-world work.

Primary source: Hugging Face · via ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

CEO Fired Developers To Make Room For AI. Developers Create Open Source AI CEO

A CEO has dismissed development staff to prioritize AI automation, leading developers to create an open-source AI CEO in response.

Pollen Robotics (Hugging Face) Microduck

Pollen Robotics introduces Microduck, a new AI-powered robot, on Hugging Face, highlighting advancements in robotics and AI integration.

The Future Of AI: How To Design A Grok Bot With Grok Bot On X.ai

xAI reveals a new project involving Grok AI in the design process of a system called Grok Bot, but details on its form, capabilities, and development stage remain unclear.

Sentiment Analysis Tools to Gauge Audience Response

Harness powerful sentiment analysis tools to gauge audience response and uncover emotional insights that can transform your engagement strategies.