DeepSeek-V4-Pro
DeepSeek's frontier MoE flagship, generally available since August 2026 with substantially stronger agentic tool use and code execution
DeepSeek-V4-Pro
DeepSeek β’ April 2026
Training Data
32+ trillion tokens, up to early 2026
DeepSeek-V4-Pro
April 2026
Parameters
1.6 trillion (49B active)
Training Method
MoE with hybrid attention (CSA + HCA), Muon optimizer, two-stage post-training
Context Window
1,000,000 tokens
Knowledge Cutoff
Not disclosed
Key Features
Hybrid Compressed Attention β’ Manifold-Constrained Hyper-Connections β’ Thinking Effort Levels (Low/High/Max) β’ Native Responses API β’ Open Weights (MIT)
Capabilities
Reasoning: Outstanding
Coding: Outstanding
Agentic: Outstanding
What's New in This Version
GA release (0813) on 13 Aug 2026 sharply raised agentic scores over the April preview: Terminal-Bench 2.1 72.1 β 87.9, DeepSWE 12.8 β 62.7, CyberGym 52.7 β 83.3, Toolathlon-Verified 55.9 β 74.1, NL2Repo 61.5. 1M context with up to 384K output tokens; 27% inference FLOPs and 10% KV cache vs V3.2 at 1M tokens
DeepSeek's frontier MoE flagship, generally available since August 2026 with substantially stronger agentic tool use and code execution
What's New in This Version
GA release (0813) on 13 Aug 2026 sharply raised agentic scores over the April preview: Terminal-Bench 2.1 72.1 β 87.9, DeepSWE 12.8 β 62.7, CyberGym 52.7 β 83.3, Toolathlon-Verified 55.9 β 74.1, NL2Repo 61.5. 1M context with up to 384K output tokens; 27% inference FLOPs and 10% KV cache vs V3.2 at 1M tokens
Technical Specifications
Key Features
Capabilities
Other DeepSeek Models
Explore more models from DeepSeek
DeepSeek-V4-Flash
DeepSeek's smaller, fast variant of V4 β same architecture at a fraction of the cost and latency
DeepSeek-V3.2
DeepSeek's latest flagship model matching GPT-5 performance with integrated tool-use thinking
DeepSeek-V3.2-Speciale
DeepSeek's competition-focused variant (EXPIRED Dec 15, 2025 - was temporary API-only release)