In November 2021, getting GPT-3 to produce a million tokens of text cost $60 through the API. Three years later, a model of at least that quality produced the same million for $0.061: a thousandfold fall, running at roughly 10x per year, faster than compute costs fell during the PC era or bandwidth during the dot-com build-out1. Stanford's AI Index clocked the GPT-3.5-class version of the same collapse at 280x in about eighteen months2. Epoch AI, tracking prices across benchmarks, found the median decline accelerating to around 200x per year since January 20243.

A comparability note belongs up front. These are price declines at matched capability, benchmark-equivalent models compared across time, not the price of the frontier model of the day, and they mix input with output tokens, quality with latency, and list prices with negotiated ones. The direction and rough magnitude survive every reasonable adjustment; the precise multiplier depends on which quality bar and which token mix you hold fixed.

At the same time, the other price in the industry moved the other way. The amortised cost of the final training run of a frontier model has grown, on Epoch's estimated trend for frontier runs, about 2.4x per year since 20164. GPT-4's run cost tens of millions of dollars, roughly $40 million in Epoch's hardware-amortised accounting, $78 million at cloud prices, over $100 million in its maker's telling4. If the trend holds, the largest runs pass $1 billion around 20274. Using intelligence deflates; making the frontier of it compounds.

$60 to $0.06

Price per million tokens for GPT-3-class output, November 2021 to November 2024. Frontier training costs grew 2.4x a year over the same period.

Andreessen Horowitz; Epoch AI

Why serving collapsed

No single trick did it; five compounding ones did1. GPU generations kept arriving. Quantisation cut the arithmetic from 16-bit to 4-bit, a factor of four by itself. Inference software learned to batch, cache, and stream around memory bottlenecks. Post-training techniques taught small models what only giant ones used to know, so a few billion parameters now outperform the 175-billion-parameter original. And open-weight releases from Meta, Mistral, and the Chinese labs turned model-serving into a commodity business where a dozen providers race each other to the marginal cost of electricity.

The last mechanism matters most for where prices settle. A proprietary model prices against its next-best rival; an open-weight model prices against anyone's cloud bill. Every time a frontier capability is replicated in open weights, typically within a year or two, the price of that capability falls to commodity serving margins, and the frontier labs must either move up or watch their margins move down. The 280x is not generosity. It is what competition does when the recipe leaks.

January 2025 supplied the canonical demonstration. DeepSeek, a Chinese lab, released near-frontier open weights alongside a paper describing a training budget far below what the market had assumed such capability required, and serving prices across the industry lurched down within weeks. Whatever the accounting behind the disclosed figure, the market lesson stood: the replication lag had shortened again, and it had shortened from a direction export controls were supposed to have closed. Price deflation in this industry is not only an efficiency story; it is also a proliferation story.

Why the frontier compounds

Training costs rise because capability still scales with compute, and compute at the frontier is bought in units of gigawatt-datacentres and six-figure GPU fleets. Epoch's decomposition attributes the 2.4x annual growth mostly to hardware quantity: bigger clusters running longer on dearer chips4. The reasoning-model turn added a twist in 2024-25: models that think longer per query, whose improvement is bought partly at inference time, which is why OpenAI's o1 launched costing the same $60 per million output tokens that GPT-3 did at its debut1. The frontier price reappears at the top of every cycle even as the floor falls out below.

The billion-dollar run has a corollary the industry says quietly: at that price, the set of institutions that can finance a frontier attempt shrinks to a handful of firms backed by the largest balance sheets in corporate history, plus whichever states decide the capability is sovereign. Training cost growth is a concentration mechanism, the same ratchet at work in chip fabrication, where each generation's capital bill prunes the field of those who can pay for the next.

Living in the gap

Every AI business model is a position on these two curves. The application companies ride the falling one: a feature that was uneconomic at $20 per million tokens is a rounding error at twenty cents, which is why capabilities move from demo to default within quarters. The lab business model must clear the rising one: frontier revenue has to cover frontier training plus the deflation of last year's product into this year's commodity. The scariest sentence in an AI lab's finance meeting is not about competitors; it is that the product being amortised depreciates at 10x a year.

The gap also explains the industry's structural bargains. Model makers sell compute commitments to clouds and take equity from them because the training bill needs a balance-sheet partner; clouds oblige because inference at commodity prices still sells electricity and racks at volume. Meanwhile buyers of AI capability face a permanently good deal getting better, with one caveat: the cheap tier is always one capability generation behind the expensive one, and the value of that gap, months of capability lead, priced at hundreds of times the commodity rate, is the actual product the frontier labs sell.

The two curves, in numbers

Measure

Value

GPT-3-class output, Nov 2021

$60 per million tokens

Same quality, Nov 2024

$0.06 per million tokens

GPT-3.5-class decline

280x in about 18 months

Median price decline since Jan 2024

About 200x per year

Frontier training cost growth

2.4x per year since 2016

Largest runs, on trend

Over $1 billion by 2027

Andreessen Horowitz; Stanford AI Index; Epoch AI

The Jevons clause

Every technology that gets radically cheaper meets the same nineteenth-century observation: efficiency grows consumption. Coal-efficient engines burned more coal in aggregate; cheap lumens lit the night; and tokens at a ten-thousandth of their 2021 price are being spent ten-thousand-fold, in agents that read whole codebases, pipelines that draft and re-draft, products that call a model where they once called a database. The deflation is not shrinking the industry's revenue pool; it is what allows the pool's expansion into tasks that were never worth $60 a million tokens and are obviously worth six cents. The falling price is the business model.

For buyers the practical rules fall out directly. Never sign long commitments at today's per-token price for tomorrow's volumes. Build so the model underneath can be swapped, because the cheapest adequate model changes quarterly. And treat the premium tier as an option on capability lead rather than a subscription to intelligence, renewed only while the lead exists for your task. Procurement departments that learned hardware's depreciation curves over decades are relearning them at software speed.

What deflation does not deflate

The falling token price hides the bill's migration, not its disappearance. Aggregate inference spend rises even as unit costs collapse, because usage grows faster than prices fall; the data-centre build-out is being financed on exactly that arithmetic. And reasoning workloads recomplicate the unit economics: a query that burns ten thousand thinking tokens at a cheap rate can cost more than one that burned a hundred at an expensive one. Cheap tokens do not guarantee cheap answers; they guarantee that expensive answers buy more thinking.

Nor does deflation reach every task equally. Epoch's data shows the fastest declines where benchmarks are saturated and open replacements exist, and much slower ones at the capability edge3. Prices are a map of competition, and competition is a map of replication lag. Where the lag is short, tokens approach free; where a single lab holds the capability, pricing power holds too, for as long as the moat does, which recent history suggests is quarters, not years.

The macro consequence is worth spelling out, because the two curves are financing each other. The hundreds of billions in datacentre capital expenditure are a bet that aggregate inference demand, riding the falling curve, grows into the capacity being built for the rising one. If token deflation keeps expanding usage faster than it erodes revenue per token, the bet pays and the build-out looks like the railways that worked. If usage saturates while the training bill compounds, the industry has built the other kind of railway. The spread between those outcomes is the largest single uncertainty in technology markets today, and it is set by the two price curves in this piece.

What to watch

Three numbers carry the story forward. The street price of matching each new frontier model, and the delay before it, because that pair sets every AI product's margin structure. The cost of the largest disclosed training run against Epoch's billion-dollar trend line4, because the first public billion-dollar run will mark the point where frontier AI's capital structure looks like aerospace. And the ratio of industry inference revenue to training capex, because that is the number that says whether the two curves ever meet in a business that compounds, or whether the gap is being financed, indefinitely, on the belief that they must. Both curves have held for years. Neither is a law of nature.

  1. Guido Appenzeller, Welcome to LLMflation, Andreessen Horowitz (November 2024): GPT-3 at $60 per million tokens in November 2021 against $0.06 for Llama 3.2 3B in November 2024; roughly 10x annual decline; mechanisms from quantisation to open-weight competition; o1 launching at $60 per million output tokens.

  2. Stanford HAI, AI Index 2025, as summarised in inference-price surveys: about a 280-fold fall for GPT-3.5-equivalent quality in roughly 18 months.

  3. Epoch AI, LLM inference prices have fallen rapidly but unequally across tasks: median decline near 50x per year across benchmarks, about 200x per year since January 2024, with wide variation by task.

  4. Epoch AI, How much does it cost to train frontier AI models? and Cottier et al., The rising costs of training frontier AI models: amortised final-run costs growing about 2.4x per year since 2016; GPT-4 near $40 million amortised, $78 million at cloud prices in the Stanford accounting, over $100 million per Sam Altman; largest runs projected past $1 billion by 2027.