← Choose models

Historical benchmark snapshot

Evaluation metricGemini 3.1 Pro ↗High · GoogleClaude Opus 4.6 ↗Max · Anthropic
SWE-bench VerifiedResolved %80.680.8
GPQA DiamondPass@1 %94.391.3
HMMT 2026 FebPass@1 %94.796.2
MMLU-ProEM %91.089.1
LiveCodeBenchPass@1 %91.788.8
Terminal-Bench 2.0Accuracy %68.565.4
SWE-bench ProResolved %54.257.3
Humanity’s Last ExamPass@1 %44.440.0
BrowseCompPass@1 %85.983.7
MRCR 1MMMR %76.392.9
Model weightsClosed weightsClosed weights
Source: DeepSeek V4 preview model card. Checked 2026-09-24. Highest reported values appear in green, but reasoning budgets and tool frameworks may differ. These scores are not an evaluation of newer models.