All posts

Deepseek V4-Pro update boosts agent scores, hikes API prices

Manaal KhanAugust 13, 2026 at 10:32 PM4 min read
Deepseek V4-Pro update boosts agent scores, hikes API prices

Deepseek's V4-Pro model gets a significant benchmark upgrade today, but developers will pay more for API access starting August 16. The new build, V4-Pro-0813, pushes Terminal Bench 2.1 scores from 72.1 to 87.9 and ships alongside Deepseek Harness, the company's agent software now released under an MIT license.

Deepseek V4-Pro update boosts agent scores, hikes API prices
Source: The Decoder

The price changes hit hardest on cache hits, which jump from $0.003625 to $0.022 per million tokens during off-peak hours. That's roughly a 6x increase. For agent workflows that repeatedly read the same files, this shift fundamentally changes the cost calculus.

Advertisements

What changed in V4-Pro-0813?

DeepSeek V4 Pro 0813 with Major Agent Upgrade: Tested Locally

The model keeps its one-million-token context window and existing parameter count. Deepseek says current integrations will continue working without modification. The company added native support for OpenAI's Responses API with Codex integration, and reasoning effort now comes in three levels: low, high, and max. Deepseek recommends the middle setting for typical agent tasks.

The benchmark jumps are striking. DeepSWE scores climbed from 12.8 to 62.7. According to Deepseek's own table, the new build beats Claude Opus 4.8 on several agent benchmarks. But context matters.

Benchmark table comparing Deepseek V4-Pro-0813 against V4-Flash, GLM-5.2, Kimi K3, Claude Opus 4.8, and Fable 5 across ten agent and coding tests
Benchmark-Tabelle vergleicht V4-Pro-0813 mit V4-Flash-0731, beiden Preview-Builds sowie GLM-5.2, Kimi-K3, Opus-4.8 und Fable 5 in zehn Agenten- und Coding-Tests.

Artificial Analysis puts V4-Pro at 53 on its Intelligence Index, up from 45. That ties GLM-5.2 but trails Muse Spark (57), Qwen 3.8 Max (58), and Kimi K3 (60). Claude Opus 5 still leads at 63 points. The new build only beats Kimi K3 and Fable 5 in individual categories, not overall.

Deepseek hasn't released weights for the new build yet. The April preview version remains on Hugging Face.

Why the rush? V4 Flash was catching up

The update responds to an awkward situation. At the end of July, Deepseek shipped update 0731 for V4 Flash, which nearly matched the Pro Preview on Artificial Analysis's Intelligence Index while costing far less. The flagship needed daylight.

Deepseek Harness: open-source agent software

Deepseek Harness v0.1 ships as a Developer Preview under MIT. The company positions it against OpenAI's Codex and Claude, pitching it as a modular alternative for building autonomous agents.

The system runs on Cordis, a new plugin architecture where every feature is swappable: tools, sandboxes, sessions, UI components. A continuous session log tracks every prompt, tool call, and result. Runs can be resumed, branched, or replayed. Minimal mode strips the stack down to shell and file editor, the setup Deepseek uses for its own benchmark runs.

The software launches via npx through a local web interface, though Deepseek warns of compatibility issues. The project lead is Cui Tianyi, who joined from quantitative trading firm Jane Street in March 2026. When the team sought beta testers in early August, 712 projects signed up within three days.

ℹ️

Logicity's Take

The Harness release is the more interesting move here. Deepseek is trying to establish an open alternative to Claude's agent stack and OpenAI Codex. The MIT license and plugin architecture could pull in developers frustrated by closed tooling. But the pricing changes complicate that pitch. If you're running agent workflows with heavy file reads, the cache hit increase will eat into any model savings. Teams building on Deepseek should model their actual usage patterns before August 16.

New API pricing: peak vs off-peak

The new rates take effect August 16 at 4:00 p.m. UTC. Deepseek first announced peak and off-peak pricing in late June but held back specific figures until now. Time-based rates have existed since February 2025, when the company discounted V3 and R1 during nighttime hours.

Deepseek V4 API pricing table showing off-peak and peak rates for V4-Flash and V4-Pro, including cache hit, cache miss, and output costs per million tokens
Preistabelle für die Deepseek-V4-API mit Off-Peak- und Peak-Sätzen für V4-Flash und V4-Pro, jeweils für Cache-Treffer, Cache-Miss und Output pro Million Token.

Peak hours run from 1 a.m. to 4 a.m. and 6 a.m. to 10 a.m. UTC, aligning with Chinese business hours. Off-peak usage costs half the peak rate. For European users, nearly the entire afternoon falls under the cheaper tier.

During off-peak hours, V4-Pro input rises from $0.435 to $0.66 per million tokens. Output jumps from $0.87 to $1.98. Peak rates double those figures to $1.32 and $3.96. The cache discount shrinks from roughly one-hundred-twentieth of regular input to one-thirtieth.

6x
Increase in cache hit pricing, the steepest jump in the new rate structure

This partially reverses the price cut Deepseek rolled out in May. Cache hits will actually cost more than they did before that reduction. The timing coincides with Deepseek raising new capital and preparing for an IPO.

ℹ️

Need Help Implementing This?

If you're building agent workflows on Deepseek's API and need help modeling the cost impact of the new pricing structure, reach out to Logicity's consulting team.

Source: The Decoder / Jonathan Kemper

M

Manaal Khan

Tech & Innovation Writer

Produced with AI assistance and reviewed by the Logicity editorial team. Learn more in our Editorial Policy.