SIDE BY SIDE
Compare selected AI models side by side.
Compare uses, prices, speed evidence, and version details for up to three models. Language, video, decision models and historical evaluations are shown separately.
Historical benchmark snapshot
| Evaluation metric | Claude Opus 4.6 ↗Max · Anthropic | GPT-5.4 ↗xHigh · OpenAI |
|---|---|---|
| SWE-bench VerifiedResolved % | 80.8 | Not reported |
| GPQA DiamondPass@1 % | 91.3 | 93.0 |
| HMMT 2026 FebPass@1 % | 96.2 | 97.7 |
| MMLU-ProEM % | 89.1 | 87.5 |
| LiveCodeBenchPass@1 % | 88.8 | Not reported |
| Terminal-Bench 2.0Accuracy % | 65.4 | 75.1 |
| SWE-bench ProResolved % | 57.3 | 57.7 |
| Humanity’s Last ExamPass@1 % | 40.0 | 39.8 |
| BrowseCompPass@1 % | 83.7 | 82.7 |
| MRCR 1MMMR % | 92.9 | Not reported |
| Model weights | Closed weights | Closed weights |
Source: DeepSeek V4 preview model card. Checked 2026-09-24. Highest reported values appear in green, but reasoning budgets and tool frameworks may differ. These scores are not an evaluation of newer models.