AI Performance Costs Are Dropping Faster Than Any Previous Technology
Tech

AI Performance Costs Are Dropping Faster Than Any Previous Technology

TechNews Editorial
TechNews EditorialSep 24, 2026 · 3 min read
Share

The price of reaching a fixed performance level on select artificial intelligence benchmarks has dropped sharply since 2023. Epoch AI tracks artificial intelligence trends. The research organization says costs are falling by about 47 percent per quarter on average. This equals about 13 times per year. The group says no other transformative technology declined that fast. The number reflects market prices for a fixed benchmark score. It does not measure pure algorithmic or architectural progress or real-life productivity costs.

MIT researchers looking at comparable data see costs dropping 5 to 10 times annually. They strip out cheaper hardware and competitive pricing pressure. After doing this, they put the actual gain in algorithmic efficiency at about 3 times per year. Peak performance per run can actually get pricier. Newer reasoning models burn through a lot more compute per task. Both studies ask different things. Matching last year capability is dramatically cheaper. Running the current best model often costs significantly more per query.

Epoch uses OpenAI model o3 as an example. In early 2025, o3 scored 75 percent on GPQA Diamond. This is a PhD-level science test. It ran at an estimated 30 cents per question. Eighteen months later, a GPT-5.6 family model hit the same score for four hundredths of a cent. Epoch says that is 1 over 725 of the original price. If cars dropped that fast, a 50,000 euro vehicle would cost less than 70 euros. OpenAI launched the even cheaper GPT-6 Sol and Luna models just days ago. The gap has likely widened further.

Epoch bases its analysis on five benchmarks spanning math, science, and logic puzzles. Because that is a narrow sample, the organization calls its findings reasonable but rough measurements based on the best available data. Algorithmic gains only explain part of the drop. Hans Gundlach and his MIT colleagues use pricing data from the comparison platform Artificial Analysis. Their data covers April 2024 through November 2025. They evaluate far more models per test than previous research. Their numbers are lower than Epoch numbers partly because newer reasoning models throw extra test-time compute at hard problems. This drives up the cost per correct answer even while per-token prices keep falling. Tokens are the text units providers bill for. Comparing them alone misses the full picture.

When MIT breaks down what pushes prices lower, cheaper hardware accounts for some of it. Competition accounts for more. The researchers control for competition by looking at open models separately. What remains is the pure algorithmic efficiency gain. This comes out to about 3 times per year. Epoch figure is higher because it does not strip out hardware and competition effects.

MIT also found that some performance gains simply come from spending more compute. A new model beats its predecessor on GPQA Diamond. It looks like progress from the outside. The authors estimate a big chunk of the improvement comes from using more processing power per question. It scores higher and costs more to run. Coding and math benchmarks show a smaller version of the same pattern. Not all benchmark progress is efficiency progress. New models roll together better training, better data, better architecture, and more test-time compute. A single score blends all of that.

There is also the benchmaxxing problem. Artificial intelligence companies could optimize for well-known tests. This would inflate scores without real-world payoff. Epoch tries to guard against this by including one test named Mystery Game Puzzles. This test is based on a game kept secret. Costs drop slowest on that test. This fits the benchmaxxing theory. It could also just reflect the task format or noise in the data.

Price alone will not tell you which model to pick. Broad cost averages do not help much when choosing a model for a specific job. Platforms like Artificial Analysis rank models across quality, price, latency, context window, and output speed. The cheapest option rarely wins on every dimension. A low-cost model with high latency is useless for a real-time chatbot. A powerful reasoning model might be too slow for automated workflows. A pricier frontier model could still save money if it gets things right more often and cuts down on retries. None of that shows up in a simple price-per-token comparison.

Newsletter

Get the best AI & tech news daily

A concise daily digest. Unsubscribe anytime.

We use your email only to send this newsletter.

Related Stories