AI · Math Benchmarks

AI Math Reasoning (Mock AIME)

Competition mathematics: from under 1% to 100% in three years.

Saturated — A hard problem became a solved problem: the frontier reached 100% in April 2026.

Epoch AI runs the frontier models against the OTIS Mock AIME set: competition problems written to match the American Invitational Mathematics Examination without reusing its questions. The mock set matters, because real AIME problems leak into training data and make published AIME scores unreliable. The frontier went from GPT-4’s 0.6% in March 2023 to GPT-5.5 Pro’s 100% in April 2026. The largest step came in September 2024, when o1-mini took the record from 9.7% to 46.9%, trained to reason step by step before answering. Reaching 100% from there took another nineteen months. Each point is the highest mean score any model had reached at that date, taken from Epoch’s published run data. At 100% the set is saturated and cannot record what a stronger model would do.

Source: Epoch AI benchmark runs (OTIS Mock AIME 2024-2025)

Download the data: ai_math.csv · ai_math.json

AI Math Reasoning (Mock AIME): every plotted point, in %
YearValueNoteProjected
March 20230.6%GPT-4 (Mar 2023)
June 20231.1%GPT-4 (Jun 2023)
July 20232.5%Claude 2
February 20244.7%Claude 3 Opus
April 20246.7%GPT-4 Turbo (Apr 2024)
May 20246.8%Gemini 1.5 Pro
July 20246.9%GPT-4o mini
July 20249.7%Llama 3.1-405B
September 202446.9%o1-mini: reasoning arrives
December 202473.3%o1
January 202576.9%o3-mini
April 202577.8%Grok-3 mini
April 202584.4%o3
July 202586.7%Qwen3-235B-A22B (Jul 2025)
August 202588.9%gpt-oss-120b
August 202591.4%GPT-5
December 202596.1%GPT-5.2
March 202697.8%GPT-5.4
April 202697.8%Claude Opus 4.7
April 2026100%GPT-5.5 Pro: saturated

All curves · The chart is drawn in your browser and needs JavaScript. Every number it plots is in the table above.