← Choose models

Historical benchmark snapshot

Evaluation metricDeepSeek V4 Flash ↗Max · DeepSeekClaude Opus 4.6 ↗Max · Anthropic
SWE-bench VerifiedResolved %79.080.8
GPQA DiamondPass@1 %88.191.3
HMMT 2026 FebPass@1 %94.896.2
MMLU-ProEM %86.289.1
LiveCodeBenchPass@1 %91.688.8
Terminal-Bench 2.0Accuracy %56.965.4
SWE-bench ProResolved %52.657.3
Humanity’s Last ExamPass@1 %34.840.0
BrowseCompPass@1 %73.283.7
MRCR 1MMMR %78.792.9
Model weightsOpen weightsClosed weights
Source: DeepSeek V4 preview model card. Checked 2026-09-24. Highest reported values appear in green, but reasoning budgets and tool frameworks may differ. These scores are not an evaluation of newer models.