AI · Research Mathematics
FrontierMath
Unpublished problems, vetted by mathematicians: eighteen of the 285 are left.
- Zero to 267 of 285 in two and a half years
- Unsolved share fell 7× in twelve months
- Epoch’s own runs, with Python and 1M tokens
- OpenAI funded it; 12 of 13 records are its own
FrontierMath is Epoch AI’s attempt to write mathematics a model cannot have read: original problems, unpublished, vetted by working mathematicians, and Epoch says a typical one costs a researcher in the right branch multiple hours, the hardest of them multiple days. This curve plots the private core of Tiers 1-3 — 285 problems — from a flat zero in January 2024. Every score is Epoch’s own run, not a lab’s announcement, and every model gets tools: a Python interpreter, a hard ceiling of a million tokens. This is mathematics with a computer to hand. The two largest jumps are behind it, GPT-5 in August 2025 and GPT-5.2 Pro in December, each worth more problems than the eighteen that remain. Read it from the other end, as the share still unsolved: that fell sevenfold in the twelve months to GPT-6 Astra’s release on 3 September 2026, halving about every 4.3 months. There is not much left to halve. Above Tiers 1-3 sits Tier 4, forty-one private research problems, each a several-week project to write, where Astra scores 97.6%. Beyond it, FrontierMath Erdős: sixty-eight conjectures Paul Erdős posed or studied, open in August 2026 and stated in Lean, so a proof has to survive a machine. Astra has two; every other model Epoch has run on that set has none. In June 2026 Epoch rebuilt the benchmark, “a major update, addressing errors in 42% of problems”, and re-ran models on the new set rather than adjusting the scores it had. So this history is younger than it looks: every run behind it dates from summer 2026, and the oldest point was measured on 27 August, two and a half years after its model shipped. Five earlier versions of the set are not on this line, including the 25.2% OpenAI announced for o3 in December 2024, on a 180-question version. Epoch’s own statement reads: “FrontierMath was developed with funding from OpenAI, who has exclusive access to a subset of the benchmark.” OpenAI commissioned the questions and holds the problem statements and solutions up to the December 2024 version, less 53 solutions withheld at random from the February 2025 one and 20 of the 50 Tier 4 problems, both kept back for holdout evaluation. Twelve of the thirteen records here are OpenAI models, partly a choice about who was re-run: on the v1 set the line was eight of twelve, and its four non-OpenAI holders have no v2 run. The holdout is the check on that, and Epoch’s run data carries no holdout task: its result is not in the file this curve reads.
| Year | Value | Note | Projected |
|---|---|---|---|
| January 2024 | 0% | GPT-3.5 Turbo: a flat zero | |
| April 2024 | 0.7% | GPT-4 Turbo (Apr 2024): 2 of 285 | |
| December 2024 | 14.74% | o1 | |
| January 2025 | 18.6% | o3-mini | |
| April 2025 | 36.14% | o4-mini | |
| August 2025 | 55.44% | GPT-5: first past half | |
| October 2025 | 55.79% | GPT-5 Pro | |
| December 2025 | 74% | GPT-5.2 Pro | |
| March 2026 | 82.46% | GPT-5.4 Pro | |
| April 2026 | 87.72% | GPT-5.5 Pro | |
| July 2026 | 89.12% | GPT-5.6 Sol | |
| September 2026 | 90.18% | Claude Fable 5.1: the only non-OpenAI record | |
| September 2026 | 93.68% | GPT-6 Astra (2026-09-03): 18 problems left |