AI Achieves Perfect Score at Math Olympiad for First Time
For the first time in history, artificial intelligence systems have achieved a perfect score on the International Mathematical Olympiad (IMO), the world’s most prestigious mathematics competition for high school students. Two Chinese technology companies — Huawei and Xiaohongshu (RedNote) — announced on July 22–23 that their respective AI models scored a flawless 42 out of 42 on the problems from the 67th IMO, held this month in Shanghai. The milestone marks a dramatic leap from 2025, when the best AI systems reached gold-medal level but fell short of perfection, and signals that machine reasoning has reached the highest tier of human mathematical ability.
The Achievement
The 67th IMO took place from July 10 to 21 at Shanghai High School, drawing 666 contestants from 117 countries. Only seven human participants achieved perfect scores. Huawei’s AI system, Celia, and Xiaohongshu’s model, dots-note-3.0, both solved all six proof-based problems, each worth seven points, according to AFP via Free Malaysia Today.
“We are delighted with this result, because achieving a perfect score at the IMO is extremely challenging,” Xiaohongshu said in a statement. “Previously, no large language model had ever achieved a perfect score under the IMO’s official judging process.” Huawei stated that Celia had demonstrated “comprehensive problem-solving capabilities” across several mathematical fields.
Crucially, the AI systems were only given access to the problems after human contestants had completed their exams, and their solutions were submitted within a specified time limit. Xiaohongshu emphasized that “during testing, any form of human intervention was strictly prohibited,” with solutions forwarded to IMO organizers for grading.
A Rapid Ascent
The achievement represents an astonishing acceleration in AI mathematical reasoning. In 2024, Google DeepMind’s AlphaProof and AlphaGeometry 2 achieved a silver-medal score of 28 out of 42, but required formal theorem-proving languages and days of computation. In 2025, both Google DeepMind’s Gemini Deep Think and an experimental OpenAI model reached gold-medal level with 35 points, using natural language proofs within the competition’s time limit — yet neither could solve all six problems.
Now, just one year later, multiple systems have closed that gap entirely. According to the official IMO website, the 2026 competition featured 55 gold medals (threshold: 29 points), 105 silver medals, and 189 bronze medals.
Beyond Two Companies
The perfect-score milestone extends beyond the two Chinese firms. Deedy Das, a partner at US-based venture capital firm Menlo Ventures, told AFP that he independently tested this year’s IMO problems on four cutting-edge AI models. All four — from OpenAI, Anthropic, the startup Axiom Math, and Moonshot AI’s Kimi K3 — scored 42 out of 42. “The frontier of AI has officially moved well past IMO math,” Das wrote on LinkedIn.
This suggests that perfect mathematical reasoning at the IMO level is now a broad capability across frontier AI systems, not an isolated achievement by a single lab. As reported by South China Morning Post, Xiaohongshu’s dots-note-3.0 is the lightest model in its dots3 family and is expected to be open-sourced, which could accelerate research across the field.
How the AI Models Solved the Problems
Xiaohongshu’s dots-note-3.0 uses a technique called “recursive self-critique” — the model reviews its own intermediate reasoning, identifies errors, and revises before producing a final answer. The system received the original LaTeX versions of the problems and generated solutions in natural language, without translating them into formal theorem-proving languages. This approach mirrors how human mathematicians work: trying approaches, testing edge cases, and rewriting arguments until a rigorous proof emerges.
Huawei’s Celia demonstrated comprehensive capabilities across algebra, geometry, number theory, and combinatorics — the four core areas tested at the IMO.
Verification Questions Remain
Despite the excitement, the results have raised important verification questions. As detailed by remio.ai analysis, key unknowns include whether the model had internet access disabled, whether human researchers selected outputs from multiple attempts, and whether independent IMO coordinators graded the solutions. Google DeepMind’s 2025 result received official certification from IMO graders; Xiaohongshu’s result is currently self-reported.
Xiaohongshu has stated that dots-note-3.0 will be open-sourced, but has not yet released model weights or evaluation code, leaving independent reproduction as a future step.
Broader Implications
The perfect IMO score carries significance far beyond a single competition. It demonstrates that AI can now generate novel mathematical proofs — not merely calculate answers — at the highest level of rigor. The techniques involved, particularly recursive self-critique and natural-language reasoning, could accelerate research in mathematics, physics, engineering, and other fields requiring rigorous logical reasoning.
The achievement also carries geopolitical weight. Chinese companies (Huawei, Xiaohongshu, Moonshot AI) are demonstrating frontier AI capabilities that rival or exceed US labs in mathematical reasoning, amid intensifying US-China competition in artificial intelligence.
What’s Next
With the IMO effectively “solved” as a benchmark, the AI research community faces new questions. Future progress will need to be measured on more challenging benchmarks — perhaps problems from professional mathematics journals, or tasks requiring months of sustained reasoning. The reliability of these systems also remains an open question: does a model achieve 42/42 consistently, or only after extensive sampling?
For now, the message is clear: the world’s hardest math competition is no longer the exclusive domain of human genius. AI has officially joined the club — and it scored a perfect paper on its first try.