← Choose models

Historical benchmark snapshot

Evaluation metricKimi K2.6 ↗Thinking · Moonshot AIClaude Opus 4.6 ↗Max · Anthropic
SWE-bench VerifiedResolved %80.280.8
GPQA DiamondPass@1 %90.591.3
HMMT 2026 FebPass@1 %92.796.2
MMLU-ProEM %87.189.1
LiveCodeBenchPass@1 %89.688.8
Terminal-Bench 2.0Accuracy %66.765.4
SWE-bench ProResolved %58.657.3
Humanity’s Last ExamPass@1 %36.440.0
BrowseCompPass@1 %83.283.7
MRCR 1MMMR %Not reported92.9
Model weightsOpen weightsClosed weights
Source: DeepSeek V4 preview model card. Checked 2026-09-24. Highest reported values appear in green, but reasoning budgets and tool frameworks may differ. These scores are not an evaluation of newer models.