AI is becoming a bargain hunter's market, with a few luxury models on top
As AI model choices expand, organizations have new opportunities to balance capability with cost. The Register examines falling prices for some AI capabilities alongside rising costs associated with longer, agentic workloads, highlighting why matching models to the work is becoming increasingly important. Connect with R.B.Hall Associates, LLC to discuss how these trends may influence your organization's technology strategy.
How are AI token prices changing?
AI token pricing has shifted significantly, and it’s reshaping how companies think about AI budgets.
On the one hand, GPT-4-class capabilities have become dramatically cheaper. According to analysis cited by AI engineer Aman Panjwani, output that cost about $20 per million tokens in late 2022 is now closer to $0.40 per million tokens—about a 55x decline in under four years.
Specific model launches have accelerated this repricing. When DeepSeek released its R1 reasoning model at around $0.55 per million input tokens and $2.19 per million output tokens, compared with OpenAI’s o1-preview at $15 input and $60 output, the market effectively saw a ~97% discount for similar reasoning use cases. That kind of gap forces everyone to rethink pricing.
At the same time, frontier models are getting more expensive:
- OpenAI’s GPT-5.5 pricing reportedly doubled to about $5 per million input tokens and $30 per million output tokens.
- Google’s Gemini Flash 3.5 arrived at roughly 3–6x the price of the model it replaced.
- Anthropic’s Claude Sonnet 5 is priced lower per token than Claude Opus 4.8, but it tends to use more tokens to reach similar results, which can offset the apparent savings.
The net effect: inference for many workloads is becoming a commodity, while cutting-edge frontier inference is turning into a premium tier. For budgeting, this means:
- You can often get solid performance at a fraction of historical costs using non-frontier or open-weight models.
- But if you rely on top-tier reasoning or complex engineering tasks, expect to pay a premium and plan for that in your cost models.
Why are some companies seeing AI costs spike?
Even though per-token prices have dropped for many models, overall AI spend is rising for a lot of organizations. The main reasons are usage patterns and pricing structures, not just raw token rates.
Ameya Kanitkar, CTO of AI measurement platform Larridin, notes that about six months ago, companies were mostly worried about relatively modest subscription fees—typically $20–$100 per month per LLM subscription. Since then, two big shifts have driven costs up:
- More complex, longer-running “agentic” tasks
Models are now capable of handling more complex workflows that run longer and consume more tokens. As teams lean into these capabilities—especially in engineering operations—Larridin has seen AI costs increase about 10x between January and now in some environments. - Move from per-seat to metered pricing
Vendors like Anthropic have shifted corporate customers from per-seat pricing to metered, usage-based pricing, and tightened the permitted uses of subsidized subscription plans. That means the more your teams use AI, the more your bill scales, often in a non-linear way.
As a result, Kanitkar is now seeing companies spend roughly 10–20% of their labor cost on tokens. For example, that’s about $2,000–$4,000 per month in AI spend for a software engineer earning $200,000 annually.
However, higher spend is not automatically translating into higher productivity:
- Among Larridin’s clients, 15–30% of AI users account for more than 50% of total AI spend.
- That heavy usage often does not correlate with proportional gains in output.
When Larridin plotted token spend against developer productivity, they found an inflection point at around 35–40% of client AI spending. Beyond that, burning more tokens didn’t improve productivity. Using that as a token limit per employee allowed some organizations to cut AI costs by about 40% without changing anything else.
In short, costs are spiking because:
- Teams are running longer, more complex AI workflows.
- Metered pricing exposes every token to the P&L.
- There’s often no guardrail on usage, so a minority of users can drive a majority of spend.
Managing this means putting in place usage caps, model selection policies, and ROI tracking rather than relying on price cuts alone.
How can we optimize AI spend without losing performance?
There are several levers you can pull to reimagine your AI cost structure without sacrificing too much performance.
1. Set data-driven token limits
Larridin’s analysis shows that when they plotted token spend against developer productivity, there was an inflection point at about 35–40% of total AI spend. Beyond that, extra token usage didn’t improve output.
By using that inflection point as a per-employee token limit, some organizations were able to cut AI costs by roughly 40% with no other changes. This suggests a practical approach:
- Measure token usage vs. output (e.g., features shipped, tickets closed).
- Identify your own “no additional benefit” threshold.
- Implement soft and then hard limits around that threshold.
2. Mix frontier and open-weight models
Open-weight models are becoming a serious alternative for many workloads. Kanitkar points out that models like Kimi 2.6/2.7 and GLM 5.2 are “almost at parity” with Anthropic’s Opus 4.7/4.8 for certain tasks, while being:
- About 10x cheaper in theory, and
- Roughly 5x cheaper in practice, once you factor in slower speed and higher token consumption.
A practical pattern many companies are adopting (about 75% now use multiple models) is:
- Use open-weight or lower-cost models for routine coding, documentation, and internal analysis.
- Reserve frontier models (like Anthropic’s Opus) for complex engineering and reasoning tasks where they clearly outperform. Larridin’s data shows enterprises still direct almost half of their AI spend to Opus for this reason.
3. Align pricing models with usage patterns
As vendors move from per-seat to metered pricing, it becomes more important to:
- Audit where AI is used (e.g., engineering ops vs. customer-facing agents).
- Identify workflows that are long-running or “agentic” and see if they can be simplified or partially automated instead of fully delegated.
- Negotiate or choose plans that match your actual usage profile, not just headline per-token rates.
4. Treat AI as a shared, measured capability
Given that 15–30% of users can drive over 50% of AI spend, it helps to:
- Make AI usage visible across teams (dashboards, reports).
- Set expectations for when to use premium models vs. commodity ones.
- Regularly review high-usage accounts to ensure spend aligns with business value.
By combining usage caps, smart model selection, and better visibility, companies can reshape their AI cost curve while still benefiting from strong performance and developer productivity.


