Token unit prices have dropped by 98%, yet AI bills have increased by 320%—Pricing design in an era where 'the cheaper it gets, the more it costs'
"AI is getting cheaper, so if you wait, your budget will be easier to manage"—if you think that, the 2026 data shows a completely different picture. Token unit prices have fallen by about 98% since 2022. Yet, corporate AI spending has increased by 320% over the same period [The Spend].
Hello, this is Mizuno. As a manager in the MA/CRM industry, I design AI businesses in my day job and support startup sales strategies as a side hustle. I am also a musician. I have previously designed service plans for AI utilization, and I have struggled with the decision of whether to go with a flat-rate or usage-based model. Today, drawing on that experience, I will write about this "the cheaper it gets, the more it costs" paradox from both the seller's and the buyer's perspectives.
※This article is based on research of public information and personal assessment. Since figures depend on the definitions of each survey, please be sure to check the primary sources when citing.
1. What is happening—the collapse of unit prices and the explosion of consumption
Let's look at the numbers. The average corporate AI budget is expected to grow from $1.2 million annually in 2024 to about $7 million by 2026 [Oplexa]. According to Ramp economists, the average company's token spending is 13 times higher compared to January 2025. And for some Fortune 500 companies, monthly inference costs are said to be reaching the tens of millions of dollars.
The culprit is the shift from chatbots to agents. While a chat where a human types a prompt might only take one round trip, an autonomous agent chains 10 to 20 LLM calls behind the scenes for a single task. Some analyses suggest that a typical agent job consumes about 96,000 tokens—the text of an entire novel—before providing an answer. A process that cost $0.04 in 2023 costs about $1.20 in a 2026 agent configuration. Even though unit prices have dropped, consumption is increasing at a pace that overwhelmingly exceeds that reduction.
The situation has reached the point where standardization bodies are taking action. In June 2026, the Linux Foundation established the "Tokenomics Foundation" to begin standardizing AI cost measurement and billing telemetry [Practical Logix].
2. As a seller—the pitfall I fell into when designing pricing plans
I have a bitter realization regarding this structure. When I designed a service plan for AI utilization, I created a plan that sold "peace of mind with a flat monthly rate." It seemed like the right answer to lower the psychological hurdles for customers.
But when you think about the backend, a flat-rate plan is designed so that the provider bears all the risk of exploding consumption. The more agent-based processing increases, the more unpredictable the costs become. Conversely, if you switch to usage-based billing, the cost risk shifts to the customer, but then you face the risk of churn due to "bill shock" on the customer side. Gartner's research also warns that many vendors do not even disclose how they calculate token consumption, making it almost impossible for companies to forecast costs [GovInfoSecurity].
Looking back now, I think the right answer is to stop worrying about the binary choice between flat-rate and usage-based, and instead 'show the consumption from the start'. Even with a flat-rate plan, disclose a consumption meter to the customer and show them "what it would have cost if it were usage-based." This builds trust and also serves as insurance for when you eventually need to revise prices. This is a lesson I want to share with everyone who sells AI services.
3. As a buyer—cost management is a design job, not a finance job
On the other hand, I am also a buyer and a user. As someone who runs agents in parallel on a daily basis, I can state with certainty that AI costs can change by an order of magnitude depending on how you design their usage.
The measures considered effective in enterprise settings are surprisingly grounded [Thoughtworks].
- Routing: Send simple tasks to smaller models, and reserve frontier models only for complex reasoning.
- Caching: Don't make the model infer the same question every time. Reuse results for semantically identical queries.
- Circuit breakers: Implement a mechanism to stop agents before they burn through tokens in an infinite loop.
In my own operations, I thoroughly differentiate usage: light models for research-related grunt work, and top-tier models only for parts that require judgment. The important thing here is that this is not the job of the finance department, but the job of the person designing the business processes. If you are someone who has designed workflows in the MA/CRM world, you can apply the concept of "which cost process to assign to which step" directly.
4. This crisis is an opportunity for those who know the front lines
When you hear "AI cost crisis," it sounds like a dark story, but I see it as an opportunity. There are two reasons.
The first is an overwhelming shortage of talent capable of cost design. While 98% of FinOps practitioners have started managing AI spending, there are still few people who can control token consumption at the operational design level. This is because it requires both knowledge of tools and knowledge of business operations.
The second is that pricing design itself becomes a product. What always comes after 'I want to introduce AI' is 'I need to do something about the AI bill.' Someone who understands both the pain of pricing for the seller and the pain of the bill for the buyer can provide advice on this. When I support sales strategies as a side job, I always make sure to include discussions about cost structure in AI-related proposals. It may not be flashy, but it is the part that is most appreciated right now.
5. Conclusion—The strategy of waiting for things to 'get cheaper' is no longer viable
To summarize:
- Token unit prices have fallen by about 98%, yet corporate AI spending has increased by 320%. The cause is an explosion in consumption due to agentification.
- Goldman Sachs predicts that token consumption will increase 24-fold by 2030 (via the aforementioned Thoughtworks/The Spend). The structure where consumption growth outpaces unit price declines will continue.
- Lesson for sellers: Instead of a binary choice between flat-rate or usage-based, a design that shows consumption from the start serves as both trust and insurance.
- Lesson for buyers: With routing, caching, and circuit breakers, costs can change by an order of magnitude depending on the design
- Cost design is a job for operational design, not finance. Those with experience in workflow design have the winning edge
'It will be cheaper if I wait a little longer' is correct regarding unit prices, but wrong regarding the bill. Only those who design how to use it can reap the benefits of cheaper AI. I am watching the meter today while assigning work to agents. Life is work, Work as Life—aggressive investment is only possible when you know the cost.
6. Reference links (Primary sources)
- The Spend: [The AI bill that nobody budgeted for]
- Oplexa: [AI Inference Cost Crisis 2026: Why Your AI Bill Is Exploding]
- Practical Logix: [AI Token Bill 2026: Inside the Enterprise FinOps Crisis]
- GovInfoSecurity: [Tokenomics: To Spend or Not to Spend]
- Thoughtworks: [Navigating today's AI token crisis]
Figures such as 98%, 320%, 13 times, and 96,000 tokens include citations via the articles above. Please check the primary reports when citing.

いいなと思ったら応援しよう!
平日はスタートアップ企業の社員、土日はたまにミュージシャン。読書や芸術、ITネタからガジェットまで興味は尽きない変人。
