Key takeaways
- The price of a token works like an interest rate for AI: it decides which uses are worth building. At a fixed level of capability it fell by 9 to 900 times a year, depending on the benchmark (Epoch AI).
- In Epoch AI’s model of a 1 GW AI data centre, 89% of the annual cost is capital being paid back. Electricity is 7%, so doubling its price raises the cost of a token by about 7% (Epoch AI, our calculation).
- In our sensitivity test on that model, writing servers and network gear off over three years instead of five raises the cost per token by about 42%. If utilisation fell from 71% to 50% and output fell with it, the cost would rise by about 39%.
- Software pushes the other way: doubling the tokens a GPU serves halves the cost of each one. Reasoning and agents spend much of the saving, and Gartner expects inference cost per agentic workflow to rise more than fivefold through 2028.
- In Anthropic’s list prices, frontier tokens get cheaper slowly: its Opus tier went from $25 to $20 per million output tokens in about a year (Anthropic, pricing page).
For a decade after 2008, the most important number in business was the price of money. With central-bank rates near zero, ideas that could not have paid a normal cost of capital suddenly could. A generation of companies was built on that arithmetic. Some became giants; others needed cheap money more than they needed customers. Oil works the same way. American shale producers told the Dallas Fed in March that they need about $66 a barrel to drill a new well profitably. Below that, the rigs go quiet and the airlines celebrate.
The token is becoming that kind of number for artificial intelligence. A token is the unit in which language models read and write, roughly three-quarters of an English word. Its price decides which uses of AI are worth trying. At a dollar per million tokens, a model can read every contract a company has ever signed. At a hundred dollars, it reads the one in dispute. Some of the most interesting ideas, such as agents that work unsupervised for hours or models that check every step of their reasoning, eat tokens the way a steel mill eats electricity. They are waiting on the price. I argued in March that intelligence is becoming a metered utility; the token is the meter.
That price has been falling faster than any central banker would dare. Epoch AI, a research group, found that the price of reaching a fixed level of capability fell by between 9 and 900 times a year, depending on the benchmark, in data running to early 2025. GPT-4-level performance on PhD-level science questions became about 40 times cheaper each year. Demand responded as the credit analogy predicts. Google says it processed 9.7 trillion tokens a month in May 2024. Two years later the figure was over 3.2 quadrillion, about 330 times more.
So what sets the price? The obvious answers are hardware, software and energy. Taken apart with numbers, two of them turn out bigger or smaller than they look, and two more drivers belong on the list.
Hardware: chips, memory and money
Start with where the capital goes. Epoch AI has modelled a hypothetical one-gigawatt AI data centre in the United States, built with Nvidia GB200 NVL72 servers. It costs about $37.9bn up front. Servers take $21.2bn, the facility $11.4bn and networking $4.9bn. Land is less than half of one percent (chart 1). An analyst estimate from BNP Paribas Equity Research, reported second hand, looks inside the servers. Of every $100 of AI capex it gives $25 to accelerators and $15 to memory. It is an illustrative split, so treat its exact figures with care.

Memory is where the hardware story turned this year. TrendForce blames persistent AI and data-centre demand for a widening gap between memory supply and demand. TrendForce reports that conventional DRAM contract prices rose by about 93 to 98% quarter on quarter in the first three months of 2026. In January, Amazon Web Services raised the price of its H200 GPU capacity blocks by about 15%. For buyers of computing power, the inputs to a token got dearer this year.
Hardware is also a financial instrument, and here the opening analogy turns literal. In Epoch’s model, 89% of the annual cost of the site is capital being paid back. The servers, network and buildings are spread over their working lives and charged a return on the money tied up in them. Epoch does not publish its rate, but its capital charges imply about 6% a year on IT equipment and about 8% on the buildings (our calculation). Raise both to 10% and a token costs about 10% more to produce (chart 2). The interest rate sits inside the token.
The assumed life of the chips matters even more. Epoch writes servers off over five years. Michael Burry, the investor, argues that hyperscalers have lengthened the useful lives of chips and servers that follow two- to three-year product cycles. Write the servers and network gear off over three years and the cost of a token rises by about 42% in Epoch’s model. How long a GPU lasts may be the most expensive accounting question in technology.

Energy is a small bill and a big bottleneck
Electricity is the driver most people overrate. In Epoch’s model, with power at 8.34 cents a kilowatt-hour, energy is $594m of an annual bill of $8.5bn, or 7%. Double the electricity price and a token costs about 7% more. An operator paying twice that rate starts at a disadvantage, though a modest one.
Energy bites hardest when it is missing. In the BNP Paribas split, grid connection and independent generation take $10 of every $100 of AI capex, more than all of the cooling. A campus waiting for its grid connection produces no tokens at all, however cheap its tariff. Energy reaches the token price mostly through capital spending and scarce connections, with the monthly bill a distant second.
Software cuts the price, then spends the savings
Software pushes the other way. Smaller models distilled from larger ones, lower-precision arithmetic and mixture-of-experts designs, which wake only part of the network for each token, all let the same silicon produce more tokens. So do better batching and caching. If software doubles the tokens a GPU can serve, the cost of each token halves. In chart 2 that is the largest swing, though we chose the size of each change. Epoch’s price series follows the cheapest model that reaches a given score, so it captures this kind of catching up, along with cheaper hardware and competition.
The catch is that software also spends what it saves. Reasoning models think before they answer, and the thinking is billed as output. Agents call a model again and again, re-sending a growing context each time. Gartner expects inference cost per agentic workflow to rise more than fivefold through 2028, even as tokens get cheaper. The price per token falls while the tokens per task climb, like cheaper petrol arriving with bigger cars. For a business, the unit that counts is the cost of a finished task.
Get the next one by email
Power, cooling, chips and the rules that shape them. No spam, unsubscribe any time.
Two more drivers: utilisation and the frontier premium
The first is utilisation. An idle GPU costs almost as much as a busy one. Epoch applies a 71% utilisation rate in its energy calculation. If we read that as the share of time the chips serve tokens, and assume output scales with it, a fall to 50% raises the cost per token by about 39% (our calculation). The same capital is simply spread over fewer tokens. Operators who keep their fleets busy can undercut those who cannot.
The second is the gap between two prices. Commodity tokens, meaning last year’s level of capability, keep getting cheaper as rivals compete. Frontier tokens move slowly. Anthropic launched Claude Opus 4.5 in November 2025 at $25 per million output tokens. Its current Opus tier lists at $20, and its top model, Fable 5.1, at $50. That is a cut of a fifth in a year, from Opus 4.5 to Opus 5.5. Epoch found a 40-fold annual fall in the price of GPT-4-level science performance. The two prices are linked the way a central bank’s policy rate is linked to what a risky borrower pays. The ambitious uses described above need the expensive kind, in bulk.
The levers compared
| Lever | In the Epoch 1 GW model | Effect on cost per token | Signal in 2026 |
|---|---|---|---|
| Server prices | $21.2bn upfront, $5.0bn a year | 30% dearer: +18% | Conventional DRAM contract prices up about 93 to 98% in Q1 (TrendForce) |
| Asset life | 5 years for IT equipment | 3 years: +42% | Burry cites two- to three-year product cycles (Techstrong) |
| Cost of capital | About 6% on IT, 8% on buildings (implied) | 10%: +10% | Not assessed here |
| Utilisation | 71%, used for energy | 50%, if output scales with it: +39% | Not assessed here |
| Electricity | 8.34 cents per kWh, 7% of annual cost | Price doubles: +7% | Not assessed here |
| Grid connection | Utility works $0.16bn | Not modelled; a missing connection means no output | BNP Paribas illustrative split: $10 per $100 of AI capex (reported) |
| Software efficiency | Not modelled | Twice the tokens per GPU: 50% less | Fixed-capability prices fell 9 to 900 times a year (Epoch AI) |
A monetary policy nobody sets
Put together, the token price behaves like a monetary policy with no central bank. Software is the dove on the committee, cutting relentlessly. DRAM shortages, scarce grid connections and the cost of capital are the hawks, and this year they have won a few votes. Commodity tokens will probably keep getting cheaper, though Epoch cautions that its fastest declines may not persist. Frontier tokens are likely to stay expensive for as long as DRAM and connected megawatts are short.
For anyone building on AI, the lesson is the one founders learned from interest rates. Know the token price your plan assumes, and how far it can rise before the plan breaks. Plenty of zero-rate companies had business models that only worked at zero. The AI businesses that last will be the ones that work at today’s frontier price and count every future cut as margin.
Sources
- Epoch AI, LLM inference price trends, 12 March 2025
- Epoch AI, AI data centre cost breakdown, May 2026
- Google, I/O 2026 keynote, 19 May 2026
- Federal Reserve Bank of Dallas, Energy Survey, 25 March 2026
- TrendForce, 1Q26 memory price outlook, 2 February 2026
- TrendForce, DRAM contract prices, 1 June 2026
- The Register, AWS raises GPU prices 15%, 5 January 2026
- Techstrong.ai, Michael Burry on hyperscaler depreciation, 11 November 2025
- Cloudnews, BNP Paribas Equity Research split of $100 of AI capex
- Gartner, inference costs per agentic workflow, 17 August 2026
- Anthropic, Introducing Claude Opus 4.5, November 2025
- Anthropic, Claude pricing
- OpenAI, reasoning models guide
- OpenAI, What are tokens
Cite this article
Meroli, S. (2026, October 11). The interest rate of intelligence: what sets the price of an AI token. ScienceShot. https://scienceshot.com/post/ai-token-price-what-sets-it
Meroli, Stefano. “The interest rate of intelligence: what sets the price of an AI token.” ScienceShot, 11 Oct. 2026, scienceshot.com/post/ai-token-price-what-sets-it.
Meroli, Stefano. “The interest rate of intelligence: what sets the price of an AI token.” ScienceShot, October 11, 2026. https://scienceshot.com/post/ai-token-price-what-sets-it.
@online{scienceshot-ai-token-price-what-2026,
author = {Meroli, Stefano},
title = {{The interest rate of intelligence: what sets the price of an AI token}},
organization = {ScienceShot},
date = {2026-10-11},
year = {2026},
url = {https://scienceshot.com/post/ai-token-price-what-sets-it}
}https://scienceshot.com/post/ai-token-price-what-sets-it
Get the weekly brief
Sourced analysis of data-centre engineering and regulation, every Monday. No spam, unsubscribe any time.






