AI · Inference Pricing
Cost of o1-Level Reasoning
Frontier reasoning cost 150× less within nineteen months of arriving.
- Cheapest model matching o1 on GPQA Diamond
- Halved ~every 12 weeks until Aug 2026
- Provider list prices, 3:1 input:output
- An upper bound: flat-rate providers only
A raw token price says little on its own, because the leading model keeps changing and so does what a token buys. This curve fixes the bar instead and lets the price move. That bar is o1, the model that made step-by-step reasoning work in December 2024, and the measure is its 76.8% on GPQA Diamond, a set of graduate-level science questions written to resist search. At each date the chart shows the cheapest model that scores at least as well, and what its maker charges. Buying o1’s reasoning cost $26.25 per million tokens on the day it launched. Six weeks later o3-mini matched it for $1.93, a thirteen-fold cut. That price stood for most of a year, until Z.ai’s GLM-4.7 undercut it at $1.00 just before Christmas 2025. March 2026 then brought two cuts in a fortnight: Google’s Gemini 3.1 Flash-Lite at $0.56 on the 3rd, and GPT-5.4 Nano at $0.46 on the 17th. DeepSeek’s V4 Flash arrived at the end of July 2026 at $0.175, a further 2.6-fold cut, and scored 91%, well clear of the bar. That is 150-fold in nineteen months, a halving just under every twelve weeks. Then the price moved the other way for the first time on this curve. Sixteen days after that record, DeepSeek switched to peak and off-peak pricing: from 16 August 2026 the same model blends to $0.33 off-peak and $0.66 at peak. The record price was real on the date it is plotted, but it is no longer offered. A record line keeps a point at the date it was observed. Under this curve’s flat-rate rule a tiered price cannot set the next point. The cheapest surviving flat rate is Z.ai’s GLM-5.3-Flash, which clears the bar at 90% on GPQA Diamond and lists at $0.15 input and $0.50 output, a blend of $0.2375. That is about half GPT-5.4 Nano’s $0.46, and still above the $0.175 record. Its 50% launch discount, which Z.ai says expires on 9 September 2026, is excluded here, because a price with a stated end date is not the flat rate this curve tracks either. Every price is the provider’s own list price, blended three parts input to one part output, which is why o1 reads $26.25 rather than its $15 input rate. Scores are Epoch AI’s own evaluations rather than self-reported figures, so the bar is measured the same way for every model. The curve tracks only providers publishing flat, tier-free rates, so the eleven-month plateau across 2025 is partly an artefact of that rule. Other models cleared the bar more cheaply during it: Moonshot’s Kimi K2 Thinking in November, and DeepSeek’s V3.2 three weeks before GLM-4.7. Their prices are either tiered by context length or no longer published now the models are superseded, and pairing a live third-party host price with a lab’s benchmark score would invent a figure neither source states. The line is an upper bound on the true floor. GPQA Diamond is also one narrow slice of what a reasoning model does.
| Year | Value | Note | Projected |
|---|---|---|---|
| December 2024 | 26.25 USD/M tokens | o1 sets the bar: $26.25 | |
| January 2025 | 1.925 USD/M tokens | o3-mini matches it for $1.93 | |
| December 2025 | 1 USD/M tokens | GLM-4.7 undercuts it: $1.00 | |
| March 2026 | 0.5625 USD/M tokens | Gemini 3.1 Flash-Lite: $0.56 | |
| March 2026 | 0.4625 USD/M tokens | GPT-5.4 Nano: $0.46 | |
| July 2026 | 0.175 USD/M tokens | DeepSeek V4 Flash |