The timing was deliberate. On Thursday DeepSeek shipped its V4-Pro flagship and, in the same breath, told developers that API rates rise between 50% and 1,100% from August 17. The scale of the increase depends on the model, token type and usage window.
V4 Pro is built for coding agents and complex autonomous tasks, using a 1.6-trillion-parameter mixture-of-experts architecture with a one-million-token context window. Initial pricing sits at $0.435 per million input tokens and $0.87 per million output tokens, with steep discounts for cache hits. The model also introduces peak and off-peak rates, a first for the company, nudging developers to shift less time-sensitive workloads into cheaper windows.
Benchmarks are strong. DeepSeek scored 87.9 on Terminal-Bench 2.1, edging past Anthropic’s Opus 4.8 and trailing Fable 5 by a whisker, and posted 83.3 on CyberGym and 62.7 on DeepSWE. The model supports OpenAI’s Responses API and offers adjustable reasoning effort, with a so-called Expert Mode in the consumer app.
The move signals a shift in strategy. DeepSeek built its global reputation on near-frontier performance at cut-rate prices, and its V4-Flash beta drew a wave of agent builders. Variable pricing acknowledges what operators of frontier models have long known: demand is uneven, and cheap off-peak compute can smooth the load.
The question now is whether developers, many of whom chose DeepSeek for cost, will absorb the increases or drift to rivals while the pricing experiment plays out.