AI · Research Mathematics

FrontierMath

Unpublished problems, vetted by mathematicians: eighteen of the 285 are left.

FrontierMath is Epoch AI’s attempt to write mathematics a model cannot have read: original problems, unpublished, vetted by working mathematicians, and Epoch says a typical one costs a researcher in the right branch multiple hours, the hardest of them multiple days. This curve plots the private core of Tiers 1-3 — 285 problems — from a flat zero in January 2024. Every score is Epoch’s own run, not a lab’s announcement, and every model gets tools: a Python interpreter, a hard ceiling of a million tokens. This is mathematics with a computer to hand. The two largest jumps are behind it, GPT-5 in August 2025 and GPT-5.2 Pro in December, each worth more problems than the eighteen that remain. Read it from the other end, as the share still unsolved: that fell sevenfold in the twelve months to GPT-6 Astra’s release on 3 September 2026, halving about every 4.3 months. There is not much left to halve. Above Tiers 1-3 sits Tier 4, forty-one private research problems, each a several-week project to write, where Astra scores 97.6%. Beyond it, FrontierMath Erdős: sixty-eight conjectures Paul Erdős posed or studied, open in August 2026 and stated in Lean, so a proof has to survive a machine. Astra has two; every other model Epoch has run on that set has none. In June 2026 Epoch rebuilt the benchmark, “a major update, addressing errors in 42% of problems”, and re-ran models on the new set rather than adjusting the scores it had. So this history is younger than it looks: every run behind it dates from summer 2026, and the oldest point was measured on 27 August, two and a half years after its model shipped. Five earlier versions of the set are not on this line, including the 25.2% OpenAI announced for o3 in December 2024, on a 180-question version. Epoch’s own statement reads: “FrontierMath was developed with funding from OpenAI, who has exclusive access to a subset of the benchmark.” OpenAI commissioned the questions and holds the problem statements and solutions up to the December 2024 version, less 53 solutions withheld at random from the February 2025 one and 20 of the 50 Tier 4 problems, both kept back for holdout evaluation. Twelve of the thirteen records here are OpenAI models, partly a choice about who was re-run: on the v1 set the line was eight of twelve, and its four non-OpenAI holders have no v2 run. The holdout is the check on that, and Epoch’s run data carries no holdout task: its result is not in the file this curve reads.

Source: Epoch AI benchmark runs (FrontierMath Tiers 1-3 v2, private)

Download the data: frontiermath.csv · frontiermath.json

FrontierMath: every plotted point, in %
YearValueNoteProjected
January 20240%GPT-3.5 Turbo: a flat zero
April 20240.7%GPT-4 Turbo (Apr 2024): 2 of 285
December 202414.74%o1
January 202518.6%o3-mini
April 202536.14%o4-mini
August 202555.44%GPT-5: first past half
October 202555.79%GPT-5 Pro
December 202574%GPT-5.2 Pro
March 202682.46%GPT-5.4 Pro
April 202687.72%GPT-5.5 Pro
July 202689.12%GPT-5.6 Sol
September 202690.18%Claude Fable 5.1: the only non-OpenAI record
September 202693.68%GPT-6 Astra (2026-09-03): 18 problems left

All curves · The chart is drawn in your browser and needs JavaScript. Every number it plots is in the table above.