First-handAIArticle··6 min read

The Price of Artificial Thought Is Falling 13x a Year, a Historic Record

Epoch AI measured how fast a given level of AI performance becomes cheap. No disruptive technology has ever gotten cheaper this quickly.

The Price of Artificial Thought Is Falling 13x a Year, a Historic Record
Source : Luke Emberson (Epoch AI) · Epoch AIView original

In brief

According to Epoch AI, the cost to reach a given performance level across five benchmarks (math, science, games) has dropped roughly 47% per quarter since 2023, a factor of 13x per year. That's far faster than electricity, batteries, compute, or DNA sequencing. The drop is even more brutal right after a performance level becomes state of the art, which says a lot about how fragile lab margins really are.

🍺 Bar-stool version

Answering a PhD-level physics multiple-choice question cost 30 cents in early 2025; eighteen months later, it costs four hundredths of a cent. It's like your brand-new car going from $50,000 to $69, and the dealer thinking that's totally normal. Electricity took eighty years to do much less, and it never learned to play chess. Except cheaper doesn't mean less spending overall: once something becomes nearly free, people end up ordering a million of it.

Key takeaways

  1. 1

    Since 2023, the cost of a fixed AI performance level has fallen by roughly 47% per quarter, or 13x per year, averaged across five benchmarks: FrontierMath (tiers 1-3), OTIS Mock AIME, GPQA Diamond, Chess Puzzles, and Mystery Game Puzzles.

  2. 2

    Math drops fastest (50–52% per quarter, 16–19x per year), game puzzles slowest (39–43%, 7–10x per year).

  3. 3

    Flagship example: 75% on GPQA Diamond cost about $0.30 per question with o3 (January 2025), then $0.0004 with GPT-5.6 Luna less than 18 months later — a 725x drop.

  4. 4

    Right after a performance level becomes state of the art, cost falls 66% per quarter (75x per year); two years later, the decline slows to 32% per quarter (4.7x per year).

  5. 5

    In log points, LLM inference has fallen 4 times faster than DNA sequencing, 6 times faster than compute, 18 times faster than lithium-ion batteries, and 54 times faster than electricity.

  6. 6

    The method reconstructs each model's full cost-performance curve from benchmark transcripts, following a procedure from the Center for AI Standards and Innovation (CAISI).

  7. 7

    The authors themselves list the limitations: benchmaxxing, the gap between benchmarks and useful work, real users who don't constantly switch models, and only three years of data.

An output price collapsing while inputs skyrocket

The AI boom is driving up the price of what it consumes: chips, electricity, and even electricians' labor. Epoch AI looks at the paradoxical flip side of this phenomenon: the price of what data centers produce is falling at an unprecedented rate.

The chosen example makes the point vividly. In late January 2025, o3 scored 75% on GPQA Diamond, a PhD-level multiple-choice test in physics, chemistry, and biology, for about 30 cents per question. Less than 18 months later, GPT-5.6 Luna gets the same score for $0.0004.

That's a 725x drop. The report puts it in a striking image: a brand-new car whose price would go from $50,000 to $69.

Measuring real cost, not token price

Previous work reasoned in terms of price per token: Guido Appenzeller (a16z) found 10x per year, an earlier Epoch analysis from March 2025 found between 9 and 900x depending on the benchmark. Reasoning models, launched with o1, broke this approach: they consume far more tokens but rely on smaller, cheaper-per-token models.

So Epoch measures the actual cost to reach a given score. Rather than running every model at every possible budget, the team adapted a CAISI method: starting from an unconstrained run, they count questions correctly solved under a token budget X, with the rest scored as random guesses. This produces a full performance-versus-spend curve.

To avoid bias, the team filled data gaps by testing models at low reasoning effort and small efficient models, including open weight models hosted on rented hardware. Validation: across five models also available via API, the gap between direct cost and API price stayed under 30%.

Statistical models presented as approximate

The object of study is the Pareto frontier: the cheapest model able to reach each performance level at a given date. A frontier that's inherently unstable, one that a single added or missing result can shift for months.

Epoch tested several approaches: smoothed frontier fitting, constrained envelope, regression across all data, stochastic frontier analysis. The chosen model is a Box-Tidwell fit of the cost frontier, which tracks the data well without producing absurd reversals. Other approaches sometimes yielded implausible results, like costs that increase.

The authors explicitly forgo confidence intervals and bootstrapping, deemed misleading in this context. The 47% figure lands right on the average computed without any model, which reinforces its credibility.

The state-of-the-art premium doesn't last

On three of the five main benchmarks (AIME, FrontierMath, GPQA Diamond), cost falls fastest right after a performance level is first achieved. On average: 66% per quarter when a level first becomes SOTA, 32% two years later.

The proposed explanation: the lab that scores a new record briefly charges a premium, before competition and technical progress crush the price. Chess puzzles and the mystery game are exceptions, showing a slight acceleration instead.

Notable detail: the fiercest competition near the state of the art mostly pits closed models against each other. Only one open weight model, QwQ-32B, appears in second place among the studied examples, even though open source pressure may play an indirect role.

A decline with no historical equivalent

Epoch considers it plausible that this pace has held since November 2021, when GPT-3's API opened commercially. A clue: GPT-3 scored 43.9 on MMLU for $60 per million tokens; Llama 2-7B scored 45.3 in July 2023 for $0.20, a 31x per year drop.

By comparison, U.S. residential electricity prices fell 1.05x per year between 1892 and 1973, lithium-ion batteries 1.16x (1991–2024), compute 1.51x (1940–2001), DNA sequencing 1.84x (2001–2025).

The limits, stated by the authors themselves

First risk: benchmaxxing, training specifically targeted at known benchmarks. Mystery Game Puzzles, whose game is kept secret to avoid this, shows a slightly slower decline (44% versus 47%), suggesting the criticism is 'valid but not fatal.'

Also, a benchmark isn't useful work, and no real user constantly switches between models to stay on the cost frontier. Three years of data is short, and averaging choices make the result vary between 42.9% and 58% per quarter.

Finally, falling prices don't mean falling spending: a task like checking thousands of scientific papers could require the equivalent of a million benchmark runs. And while passing a first-grade math test has become free, proving hard theorems has not.

That is a 725-fold drop in the price of thought in under 18 months.
The price of passing a first-grade math test may now be trivial. The price of proving hard theorems is not.
Supra-normal profits in LLM service provision may be fleeting for any given model.

Why it matters

This report puts a solid order of magnitude on a diffuse intuition: intelligence at a given level is becoming a commodity at an unprecedented pace. Two consequences deserve attention. First, the profitability window for a frontier model closes within a few quarters, which explains the frantic race between labs and raises questions about business models built on premium margins. Second, falling unit prices say nothing about the total bill: if usage explodes (agents, long reasoning, mass processing), compute demand can keep growing, which reconciles these numbers with massive data center investments. Epoch's methodological honesty deserves credit, but the '13x per year' figure should be read as a measure of what's possible for a perfectly optimizing buyer on academic benchmarks, not as the cost decline experienced by a company deploying a coding assistant. SWE-bench Verified, closer to real work, only declines by 27.5% per quarter.

#ai#llm#costs#inference#benchmarks#epoch ai
Original source
The plunging price of thought
Luke Emberson (Epoch AI)
Open the article

For you

Put it to work on your sources.

Free: this week's articles and three sources of your own. Pro: the whole archive and your sources, from €8/month.

For your team

The same machine, on your topics.

A space in your colours, your watch angles, your curators. Pilot open to three companies.

Read next