Developers can now reach DeepSeek’s V4-Flash model through a public beta of its formal API, a release tuned for AI agent workloads. The company points to scores of 82.7 on Terminal Bench 2.1 and 54.4 on DeepSWE as evidence the model handles real computer-use and software engineering tasks.
Under the hood, V4-Flash-0731 keeps the preview’s architecture and footprint while getting a retrain. The API now supports the Responses spec and is adapted for Codex, so builders can drop the model into agentic coding and automation stacks more easily.
The update applies only to the V4-Flash API. The V4-Pro API and the models powering DeepSeek’s own app and website are unchanged, the company said.
The move intensifies competition in the already crowded market for fast, cheap frontier-adjacent models. DeepSeek has positioned V4-Flash as a high-throughput workhorse for developers who want agent capabilities without the cost of its flagship models, and the public beta formalizes a release that had previously been limited.
It also lands as Chinese AI labs step up their international developer push, offering APIs that undercut US rivals on price. For startups building agent products, the choice between US and Chinese model providers has become a straightforward cost-and-capability calculation, and DeepSeek is making sure its flash tier stays in the conversation.