Deepseek's V4-Pro model gets a significant benchmark upgrade today, but developers will pay more for API access starting August 16. The new build, V4-Pro-0813, pushes Terminal Bench 2.1 scores from 72.1 to 87.9 and ships alongside Deepseek Harness, the company's agent software now released under an MIT license.

The price changes hit hardest on cache hits, which jump from $0.003625 to $0.022 per million tokens during off-peak hours. That's roughly a 6x increase. For agent workflows that repeatedly read the same files, this shift fundamentally changes the cost calculus.
What changed in V4-Pro-0813?
DeepSeek V4 Pro 0813 with Major Agent Upgrade: Tested Locally
The model keeps its one-million-token context window and existing parameter count. Deepseek says current integrations will continue working without modification. The company added native support for OpenAI's Responses API with Codex integration, and reasoning effort now comes in three levels: low, high, and max. Deepseek recommends the middle setting for typical agent tasks.
The benchmark jumps are striking. DeepSWE scores climbed from 12.8 to 62.7. According to Deepseek's own table, the new build beats Claude Opus 4.8 on several agent benchmarks. But context matters.

Artificial Analysis puts V4-Pro at 53 on its Intelligence Index, up from 45. That ties GLM-5.2 but trails Muse Spark (57), Qwen 3.8 Max (58), and Kimi K3 (60). Claude Opus 5 still leads at 63 points. The new build only beats Kimi K3 and Fable 5 in individual categories, not overall.
Deepseek hasn't released weights for the new build yet. The April preview version remains on Hugging Face.
Why the rush? V4 Flash was catching up
The update responds to an awkward situation. At the end of July, Deepseek shipped update 0731 for V4 Flash, which nearly matched the Pro Preview on Artificial Analysis's Intelligence Index while costing far less. The flagship needed daylight.
Deepseek Harness: open-source agent software
Deepseek Harness v0.1 ships as a Developer Preview under MIT. The company positions it against OpenAI's Codex and Claude, pitching it as a modular alternative for building autonomous agents.
The system runs on Cordis, a new plugin architecture where every feature is swappable: tools, sandboxes, sessions, UI components. A continuous session log tracks every prompt, tool call, and result. Runs can be resumed, branched, or replayed. Minimal mode strips the stack down to shell and file editor, the setup Deepseek uses for its own benchmark runs.
The software launches via npx through a local web interface, though Deepseek warns of compatibility issues. The project lead is Cui Tianyi, who joined from quantitative trading firm Jane Street in March 2026. When the team sought beta testers in early August, 712 projects signed up within three days.
Logicity's Take
The Harness release is the more interesting move here. Deepseek is trying to establish an open alternative to Claude's agent stack and OpenAI Codex. The MIT license and plugin architecture could pull in developers frustrated by closed tooling. But the pricing changes complicate that pitch. If you're running agent workflows with heavy file reads, the cache hit increase will eat into any model savings. Teams building on Deepseek should model their actual usage patterns before August 16.
New API pricing: peak vs off-peak
The new rates take effect August 16 at 4:00 p.m. UTC. Deepseek first announced peak and off-peak pricing in late June but held back specific figures until now. Time-based rates have existed since February 2025, when the company discounted V3 and R1 during nighttime hours.

Peak hours run from 1 a.m. to 4 a.m. and 6 a.m. to 10 a.m. UTC, aligning with Chinese business hours. Off-peak usage costs half the peak rate. For European users, nearly the entire afternoon falls under the cheaper tier.
During off-peak hours, V4-Pro input rises from $0.435 to $0.66 per million tokens. Output jumps from $0.87 to $1.98. Peak rates double those figures to $1.32 and $3.96. The cache discount shrinks from roughly one-hundred-twentieth of regular input to one-thirtieth.
This partially reverses the price cut Deepseek rolled out in May. Cache hits will actually cost more than they did before that reduction. The timing coincides with Deepseek raising new capital and preparing for an IPO.
Need Help Implementing This?
If you're building agent workflows on Deepseek's API and need help modeling the cost impact of the new pricing structure, reach out to Logicity's consulting team.
Source: The Decoder / Jonathan Kemper
Manaal Khan
Tech & Innovation Writer
Produced with AI assistance and reviewed by the Logicity editorial team. Learn more in our Editorial Policy.
Related Articles
More in AI & Machine Learning
Bezos AI Lab Gets $10B: What Project Prometheus Means
Jeff Bezos is closing a $10 billion funding round for Project Prometheus, an AI lab focused on physics-based AI for manufacturing and engineering. With a $38 billion valuation and backing from JPMorgan and BlackRock, this signals a major shift in enterprise AI investment toward industrial applications.

Kimi K2.6 Open-Weight AI: 300 Agents at a Fraction of the Cost
Moonshot AI's Kimi K2.6 matches GPT-5.4 and Claude Opus 4.6 on coding benchmarks while running 300 parallel agents. For businesses locked into expensive API contracts, this open-weight model could slash AI infrastructure costs while delivering enterprise-grade automation.




