vals.ai ran the same benchmark task across the field this month and came back with a spread that looks less like competition and more like a fire sale: GPT-5.6 Luna solved it for $0.21, GPT-5.6 Sol for $1.15, Claude Fable 5 for $2.05, all on the same harness. Kimi K3 undercuts everyone on raw token price at $3 per million input tokens against Sol's $5 and Fable 5's $10. None of that is happening in a vacuum. On our own funding tracker, we've logged more than $13 billion into AI infrastructure and compute across 17 rounds in the past two weeks alone, capital that has to be recouped somehow while the price of the thing it's building keeps falling.

We asked people who actually price or build on these models, an infrastructure engineer, a solo developer who bills clients for inference spend, a security executive whose team runs agentic workloads, and a founder who buys enough tokens to feel every price move personally, what the falling numbers actually mean. Four of them, independently, landed on the same objection before we could even ask a follow-up: the number on the pricing page is not the number that matters.

RelatedAnthropic Ships Claude Opus 5 at Half of Fable 5’s Price

Is This Price Crash Temporary or a Real Structural Shift?

Shameer Erakkath, who has spent close to 21 years designing and scaling infrastructure at Hewlett Packard Enterprise, thinks the industry has already moved past treating this as a temporary land grab, because the mechanics underneath it are real engineering gains, not just discounting. "The industry doesn't fancy on the token based tasks anymore, as the algorithmic efficiency in frontier reasoning architecture are improving a lot," he said. "The major changes coming in squeezing the models into lower precision formats like FP8 is freeing up lot of High Band Width memory and that is allowing users to share same hardware. And dynamic shortening of chain of thoughts in large reasoning models like Anthropic's Claude may drive the prices further down." That is a structural explanation, not a promotional one: FP8 quantization genuinely cuts memory footprint, shared hardware genuinely lowers the per-customer cost of serving a model, and a reasoning model that learns to think in fewer tokens genuinely costs less to run. None of it depends on a lab eating losses to win share.

Sébastien De Bollivier, who has spent two years routing production inference for SME clients across OpenAI, Anthropic, and open-source models, holds both explanations at once rather than picking a side. "It's both, and that's the tension," he said. "The $13B pouring into AI infra has to justify itself somehow. Right now, providers are buying market share. But the compression is also real, open-weight models are genuinely closing the gap on hosted APIs, which puts a structural floor on how high margins can realistically go." His read on who survives that floor is blunt. "The real question for infra investors isn't whether margins compress, they will. It's whether the companies they've backed have a moat that survives $0.05-per-task inference. Most don't."

Why Do Four Different People Say the Same Thing About Cost Per Token?

Ask what metric actually matters and the convergence is almost total. Edward Tian, co-founder of GPTZero, put it plainly: "The main question for the companies is not how much they pay per token. It is how much they pay to accomplish a task. A cheaper solution may turn out to be expensive if it leads to a bigger number of errors, requiring more retrials or making the employees spend more time checking the results." De Bollivier reaches the identical conclusion from the vendor-switching side of the business. "Frankly, 'cheapest per token' is the wrong metric for teams actually shipping products," he said. "What matters is cost-per-task-completed at acceptable quality. A model that costs $0.21 per benchmark task but requires three retries or a heavier prompt engineering layer to get reliable output isn't cheaper, it's just cheaper on paper. I've seen this pattern repeatedly when switching models for clients: the headline price drops, the hidden orchestration costs climb."

Fergal Glynn, AI security advocate and chief marketing officer at Mindgard, frames the same warning around where the savings actually leak out. "AI prices are falling rapidly, but the real story isn't about cheaper tokens, it's about reducing the cost of completed work," he said. "And this matters because retries, tool calls, and failed outputs can quickly wipe out the savings from a lower list price. One may find a model that appears cheap on paper, but still become expensive due to the looping or using more context. This is common with larger workflows." Sherif Higazy, founder and CEO of Megaton AI, who evaluates video models and burns through inference at scale doing it, draws the line in almost the same words, with one addition: the task has to earn its keep in the first place. "Cost per task is more useful than cost per token, but only insofar as the task contributes to a measurable business or revenue objective," he said. "If the value produced does not eventually exceed the inference spend, the equation does not balance and the spending cannot be justified, regardless of how cheap the tokens or individual tasks become."

Sticker price versus what a workflow actually costs after retries Two bar comparisons. The first model has a low sticker price per task, but retries, tool calls and failed outputs stack additional cost on top, bringing its real total close to a second model that costs more per task up front but needs no rework. The cheaper-looking option is not necessarily the cheaper one once the workflow is finished. STICKER PRICE VS. WHAT THE WORKFLOW ACTUALLY COSTS "CHEAP" MODEL RELIABLE MODEL Sticker: $0.21/task + retries, tool calls, failed output cleanup (hidden, not on the pricing page) List: $1.15/task no rework layer needed Real total: close to, or above, the "pricier" model genztech.blog
Fig 1 Four sources independently described this same gap between the list price and the workflow's real cost.

What Does This Mean for the Money Pouring Into AI Infrastructure?

Higazy is watching the spending side of his own company shift in real time, and it argues against the idea that falling prices mean falling demand. "We've been spending less on off-the-shelf software because we increasingly build our own internal tools. We also spend less on freelance video and image production for our materials," he said. "A growing portion of that budget now goes to inference providers instead. At the same time, our overall inference consumption is increasing rapidly, so I think that likely makes a return plausible, provided the original infrastructure economics and valuations were sensible." That is money moving from one line item to another inside a single company's budget, not new money materializing, and it is a more concrete answer than most infrastructure bulls give for where the $13 billion actually gets earned back.

RelatedGPT-5.6 Sol's #1 coding score is basically a tie

Glynn's read on the investment side is more cautious. "Above all, billions are being invested into AI infrastructure, putting more pressure on profit margins, especially if pricing keeps dropping this quickly," he said. De Bollivier's answer to which infrastructure bets survive that pressure comes back to the moat question he raised earlier: providers betting on lock-in through tooling, fine-tuning, and ecosystem depth have a credible story, because pure inference commoditizes fast once open-weight models close the gap on hosted APIs. A company whose entire pitch is "we resell tokens cheaper" doesn't have that story, no matter how good this quarter's price chart looks.

The numbers behind the price war
  • $0.21 to $2.05. What vals.ai measured for the same benchmark task across GPT-5.6 Luna, GPT-5.6 Sol, and Claude Fable 5, on identical evaluation conditions.
  • $3 to $10 per million tokens. Kimi K3's input-token price against Sol and Fable 5, the widest gap in the current top three.
  • $13 billion. Capital logged into AI infrastructure and compute across 17 rounds on our funding tracker in the two weeks before this piece.
  • $0.05 per task. The floor De Bollivier says most infrastructure bets can't survive once open-weight competition fully closes the gap.

Our Take

Nobody we talked to disputes that the sticker prices are really falling, or that the engineering behind the fall, FP8 quantization, shared hardware, shorter reasoning chains, is real rather than a subsidy dressed up as progress. What four of five sources converged on, without being asked to agree with each other, is that the sticker price was never the number worth watching in the first place. A team that switches to the cheapest model on the leaderboard and then eats the cost of retries, tool-call loops, and manual cleanup hasn't found a bargain, it has moved the same spend to a line item that doesn't show up on the pricing page. The price war is real. So is the fact that most teams are still keeping score with the wrong number.

Sources & further reading

Quotes gathered directly by GENZ TECH from sources who volunteered to comment on this story, with full attribution as agreed with each. Pricing figures per vals.ai and each provider's published rates as of late July 2026.