AI · Math Benchmarks
AI Math Reasoning (Mock AIME)
Competition mathematics: from under 1% to 100% in three years.
Saturated — A hard problem became a solved problem: the frontier reached 100% in April 2026.
- GPT-4’s 0.6% → GPT-5.5 Pro’s 100%
- o1-mini was the reasoning inflection
- Mock set: real AIME leaks into training
- Frontier = best mean score to date
Epoch AI runs the frontier models against the OTIS Mock AIME set: competition problems written to match the American Invitational Mathematics Examination without reusing its questions. The mock set matters, because real AIME problems leak into training data and make published AIME scores unreliable. The frontier went from GPT-4’s 0.6% in March 2023 to GPT-5.5 Pro’s 100% in April 2026. The largest step came in September 2024, when o1-mini took the record from 9.7% to 46.9%, trained to reason step by step before answering. Reaching 100% from there took another nineteen months. Each point is the highest mean score any model had reached at that date, taken from Epoch’s published run data. At 100% the set is saturated and cannot record what a stronger model would do.
| Year | Value | Note | Projected |
|---|---|---|---|
| March 2023 | 0.6% | GPT-4 (Mar 2023) | |
| June 2023 | 1.1% | GPT-4 (Jun 2023) | |
| July 2023 | 2.5% | Claude 2 | |
| February 2024 | 4.7% | Claude 3 Opus | |
| April 2024 | 6.7% | GPT-4 Turbo (Apr 2024) | |
| May 2024 | 6.8% | Gemini 1.5 Pro | |
| July 2024 | 6.9% | GPT-4o mini | |
| July 2024 | 9.7% | Llama 3.1-405B | |
| September 2024 | 46.9% | o1-mini: reasoning arrives | |
| December 2024 | 73.3% | o1 | |
| January 2025 | 76.9% | o3-mini | |
| April 2025 | 77.8% | Grok-3 mini | |
| April 2025 | 84.4% | o3 | |
| July 2025 | 86.7% | Qwen3-235B-A22B (Jul 2025) | |
| August 2025 | 88.9% | gpt-oss-120b | |
| August 2025 | 91.4% | GPT-5 | |
| December 2025 | 96.1% | GPT-5.2 | |
| March 2026 | 97.8% | GPT-5.4 | |
| April 2026 | 97.8% | Claude Opus 4.7 | |
| April 2026 | 100% | GPT-5.5 Pro: saturated |