AI · Inference Pricing

Cost of o1-Level Reasoning

Frontier reasoning cost 150× less within nineteen months of arriving.

A raw token price says little on its own, because the leading model keeps changing and so does what a token buys. This curve fixes the bar instead and lets the price move. That bar is o1, the model that made step-by-step reasoning work in December 2024, and the measure is its 76.8% on GPQA Diamond, a set of graduate-level science questions written to resist search. At each date the chart shows the cheapest model that scores at least as well, and what its maker charges. Buying o1’s reasoning cost $26.25 per million tokens on the day it launched. Six weeks later o3-mini matched it for $1.93, a thirteen-fold cut. That price stood for most of a year, until Z.ai’s GLM-4.7 undercut it at $1.00 just before Christmas 2025. March 2026 then brought two cuts in a fortnight: Google’s Gemini 3.1 Flash-Lite at $0.56 on the 3rd, and GPT-5.4 Nano at $0.46 on the 17th. DeepSeek’s V4 Flash arrived at the end of July 2026 at $0.175, a further 2.6-fold cut, and scored 91%, well clear of the bar. That is 150-fold in nineteen months, a halving just under every twelve weeks. Then the price moved the other way for the first time on this curve. Sixteen days after that record, DeepSeek switched to peak and off-peak pricing: from 16 August 2026 the same model blends to $0.33 off-peak and $0.66 at peak. The record price was real on the date it is plotted, but it is no longer offered. A record line keeps a point at the date it was observed. Under this curve’s flat-rate rule a tiered price cannot set the next point. The cheapest surviving flat rate is Z.ai’s GLM-5.3-Flash, which clears the bar at 90% on GPQA Diamond and lists at $0.15 input and $0.50 output, a blend of $0.2375. That is about half GPT-5.4 Nano’s $0.46, and still above the $0.175 record. Its 50% launch discount, which Z.ai says expires on 9 September 2026, is excluded here, because a price with a stated end date is not the flat rate this curve tracks either. Every price is the provider’s own list price, blended three parts input to one part output, which is why o1 reads $26.25 rather than its $15 input rate. Scores are Epoch AI’s own evaluations rather than self-reported figures, so the bar is measured the same way for every model. The curve tracks only providers publishing flat, tier-free rates, so the eleven-month plateau across 2025 is partly an artefact of that rule. Other models cleared the bar more cheaply during it: Moonshot’s Kimi K2 Thinking in November, and DeepSeek’s V3.2 three weeks before GLM-4.7. Their prices are either tiered by context length or no longer published now the models are superseded, and pairing a live third-party host price with a lab’s benchmark score would invent a figure neither source states. The line is an upper bound on the true floor. GPQA Diamond is also one narrow slice of what a reasoning model does.

Source: Epoch AI benchmarks; provider list prices

Download the data: ai_token_cost.csv · ai_token_cost.json

Cost of o1-Level Reasoning: every plotted point, in USD/M tokens
YearValueNoteProjected
December 202426.25 USD/M tokenso1 sets the bar: $26.25
January 20251.925 USD/M tokenso3-mini matches it for $1.93
December 20251 USD/M tokensGLM-4.7 undercuts it: $1.00
March 20260.5625 USD/M tokensGemini 3.1 Flash-Lite: $0.56
March 20260.4625 USD/M tokensGPT-5.4 Nano: $0.46
July 20260.175 USD/M tokensDeepSeek V4 Flash

All curves · The chart is drawn in your browser and needs JavaScript. Every number it plots is in the table above.