DeepSeek raises some V4 prices by more than 10x as AI demand strains capacity

5gDedicated

One of AI vendor DeepSeek’s biggest selling points has been its ultra-low price point, but that party’s about to end.

The Chinese model provider is raising API pricing for its V4 model family by notable margins, in some cases by more than 1,100%. The increases may not be that dramatic for all, though; the company is encouraging “more flexible workload scheduling,” with peak rates and half-price off-peak rates.

The news was tucked into the announcement of the general availability (GA) of DeepSeek V4-Pro and upgrades to VR-Flash. The new pricing takes effect for most parts of the world on August 16.

“On paper, at peak, against the right comparator, DeepSeek’s price advantage does disappear, and in places inverts,” said Sanchit Vir Gogia, chief analyst at Greyhound Research. But in practice, “the schedule’s own clock and cache hand most of it back to any buyer paying attention.”

How Flash and Pro compare now

The new API pricing structure is as follows:

Flash is now $0.22 per million input tokens (cache miss) and $0.66 per million output tokens off-peak; and $0.44 per million input tokens (cache miss) and $1.32 per million output tokens at peak.This is up from the flat rate of $0.14 for inputs (cache miss), representing a 57% to 214% increase, and $0.28 per million tokens for outputs, a 136% to 371% increase.

 Pro is now $0.66 per million input tokens (cache miss) and $1.98 per million output tokens off-peak; and $1.32 per million input tokens (cache miss) and $3.96 per million output tokens at peak.This represents an input increase of between 51% and 203% (up from $0.435) and output increase between 127% and 355% (up from $0.87).

Inputs with cache hits, when apps reuse stored prompts rather than processing similar requests from scratch, have even more dramatic pricing increases of 52% to 1,100%.

Mark Tauschek, VP of research fellowships and distinguished analyst at Info-Tech Research Group, pointed out that the increase does eliminate the price advantage that 4.0 Flash has over OpenAI 5.6 Luna at peak pricing, but not at off-peak pricing, as OpenAI has dropped Luna API pricing by 80%, off-peak.

It also doesn’t eliminate Deepseek 4.0 Pro’s price advantage over Terra, OpenAI’s GPT-5.6 mid-tier reasoning model, even at peak pricing, nor its advantage over GPT-5.6 Sol released in July, Tauschek said.

Greyhound Research’s Gogia noted that, off-peak, V4 Flash is “marginally more expensive” on input and 45% cheaper on output than Luna. Pro at peak, meanwhile, runs close to 5x Luna’s price on a representative coding-agent workload.

DeepSeek’s roughly 98% cache-hit discount, against an industry norm nearer to 90%, is the mechanism that has kept its measured cost per task at about 60% below Luna, even after Luna’s cost cut, he said.

“The schedule re-prices exactly that mechanism,” Gogia said. Flash’s edge over Luna decreases from roughly sevenfold to threefold off-peak, and 1.4 times at peak. “The cache is where the advantage genuinely erodes.”

Encouraging users to rethink their schedules

DeepSeek’s V4-Pro is now generally available, and V4-Flash is in beta. Both models have new flexible reasoning capabilities (low, high, max) and ‘thinking modes’ that use chain-of-thought (CoT) reasoning to improve answer accuracy. V4 Pro is now available on app, web, and via API, and users can try it using “Expert Mode.” V4 Flash is now in beta.

The general availability “completes a two-tier structure in which Flash serves volume and Pro is priced for complexity,” Gogia noted.

DeepSeek’s peak/off-peak pricing is a means to “allocate resources more reasonably,” the company said, to encourage users to “schedule their tasks based on actual usage.”

Gogia pointed out that with the new model, 17 of every 24 hours stay at half price, so timing becomes an economic variable, and work that can wait moves into the cheap hours. In fact, the new pricing schedule hits DeepSeek’s home market hardest and its export market lightest; Western buyers largely pay the off-peak rates.

“Usage is following economics at least as much as capability, and economics can change by schedule,” Gogia noted.

Simple supply and demand

Reading between the lines provides a more nuanced picture, Tauschek noted. “While it’s alarming to see the headlines saying DeepSeek is raising API pricing by 50%-1100%, it doesn’t really tell the whole story.”

Part of that story is demand, which is increasing exponentially. DeepSeek can’t keep up with compute requirements, and Anthropic also had a price increase for the same reason in April. And, while third-party providers have not yet reflected that trend, they’ll eventually have to, Tauschek said.

“This isn’t unexpected at all,” he noted. “It’s simple supply and demand: when demand goes up, pricing goes up, because supply becomes constrained.”

For enterprises that do use DeepSeek (many in the US do not, or can not), the new pricing is not likely to change anything, he said. Cost increases will mostly impact developers, but it will still be less expensive than most alternatives.

He pointed out that enterprises are adapting to model routing, which is critical for developers using agentic workloads. Just a few months ago, organizations were paying per-seat pricing and running up usage as a matter of course, but the market move to usage-based pricing has resulted in sticker shock akin to that of the early cloud days.

“Pricing will continue to be a big deal because CFOs are starting to ask what they’re getting for the massive AI spend,” Tauschek said.

DeepSeek pricing doesn’t change the need for compatibility, multi-modality

CIOs should read the schedule with “relief and unease,” Gogia noted. Relief because the bill is largely schedulable; unease because “a supplier that has learned to price the clock has learned something about its own leverage.”

Going forward, he predicted, Flash keeps the volume usage, Pro handles complexity, and interface compatibility lowers the cost of adoption and departure. The real question becomes whether lower economic floors, open weights, and compatible interfaces, when taken together with multi-model routing, make foundation model intelligence materially easier to substitute.

Capable inference can be produced “far below the price structures that once surrounded frontier AI,” Gogia noted, and open weights mean model developers become one of just several parties able to serve inference requirements. “The traditional software dependency changes shape when that happens,” he said.

The vendor still matters, as do capability and support, but once a workload can move between providers, and enterprises manage their own orchestration and governance, the vendor no longer owns the whole dependency, Gogia said.

The most lasting effect of DeepSeek is unlikely to be that it stayed cheapest, he noted. “It is that every provider must now explain why intelligence should command a premium once near-equivalent capability is available through several technical and commercial routes.”

This article originally appeared on InfoWorld.ComputerworldRead More