Claude Opus 5.5
Long-running coding agents and knowledge work.
Intelligence Index v4.3.2: 58 · Claude Opus 5.5 · max, default fallback
Claude API · standard processing
COMPARE MODELS
Explore language, video, and decision models by task and evidence. Check the version, serving conditions, and official sources before you choose.
Language model comparisons use the uncached input price in USD per 1 million tokens. Output is billed separately and shown on each card. Excludes tools, cache writes, storage, and tax.
TASK LEVEL × INPUT PRICE
Each model keeps its exact score height. The right side groups those scores into complex planning, everyday coding, and simple tasks.
Dot height is the exact score; the right-side regions show its task tier. Labels move downward only to avoid overlap. Select a point for details.
Not plotted because a verified index score or a positive input price is unavailable: MiniMax M2.7 Highspeed, Qwen3.8-Max.
The vertical axis groups models into three score-based rows: above 50 for complex planning, 39–50 for everyday coding, and below 39 for simple tasks. MiMo V2.6 Pro is set conservatively to 38. Scores do not replace testing on your own workload. Task buttons highlight one row without changing model placement.
“Lower cost” describes token prices, not proven value per successful task. For real value, compare correctness, retries, tools, and total billed tokens on the same tasks.
The horizontal axis uses a logarithmic scale: equal spacing means equal price ratios. Prices are USD per one million native tokens, using uncached input under each entry’s stated API conditions. Price filters, sorting, and chart positions use the same input rate. Capability evidence and pricing sources are linked in each detail page.
Model detail pages show a (fast) badge only when a published median output speed exceeds 150 tokens/s in the tested settings. That speed does not include time to first token or total task time. Speed methodology ↗
Long-running coding agents and knowledge work.
Intelligence Index v4.3.2: 58 · Claude Opus 5.5 · max, default fallback
Claude API · standard processing
General coding, visual tasks, and document workflows.
Intelligence Index v4.3.2: 56 · Claude Sonnet 5.5 · max, default fallback
Claude API · standard processing
Multimodal reasoning and tool workflows with peak and off-peak API rates.
Intelligence Index v4.3.2: 39 · DeepSeek V4.1 Flash · max
DeepSeek API · peak hours
Software engineering, autonomous agents, and enterprise workflows.
Intelligence Index v4.3.2: 41 · Gemini 3.8 Flash · high
Gemini Developer API · Standard · introductory
Text-only reasoning, complex coding, and long-running agents.
Intelligence Index v4.3.2: 45 · GLM-5.3 · max
Z.ai API · listed USD rates
Complex reasoning, software engineering, research, and document creation.
Intelligence Index v4.3.2: 53 · GPT-6 Astra · max
OpenAI API · Standard · short context
Focused, high-volume tasks with low token prices.
Intelligence Index v4.3.2: 38 · GPT-6 Luna · max
OpenAI API · Standard · short context
Coding and configurable reasoning with agent tool support.
Intelligence Index v4.3.2: 46 · Grok 4.7 · xhigh
Grok API · listed token rates
Long-horizon software engineering, reasoning, and visual knowledge work.
Intelligence Index v4.3.2: 44 · Kimi K3 · max
Kimi API · listed USD rates
Lower-cost multimodal reasoning, coding, and agent workflows; verify demanding plans before use.
Model score: 38 · conservative editorial adjustment
Xiaomi MiMo API · overseas USD · real-time
A faster serving option for M2.7 coding and agent workflows.
Intelligence Index: no reviewed score for this version
MiniMax API · Highspeed
Multimodal coding and agent tasks with a million-token context.
Intelligence Index v4.3.2: 29 · MiniMax M3 · reasoning
MiniMax API · Standard · input ≤512K
Coding, office documents, and visual understanding across long inputs.
Intelligence Index: no reviewed score for this version
Model Studio · Singapore / International
Use-case tags are editorial groupings based on official capabilities. Language models earn (fast) only with a published median output speed above 150 tokens/s in the linked test settings. Video speed labels remain provider-reported. Unknown values are excluded from price bands and sorted last.
Original research and practical experience, with short notes to help you choose.
A comparison of model–harness combinations that puts task success and cost next to each other.
How domain-specific benchmark slices and weights change the meaning of a model capability score.
A dated evaluation of GPT-6 Astra across capability, reasoning settings, token use, and task cost.
Why user preference alone does not settle factual accuracy, and how Arena adds factuality signals.
| # | Model / reasoning mode | Compare | ||||
|---|---|---|---|---|---|---|
| 01 | ✦Gemini 3.1 ProGoogle · High | 94.3 | 80.6 | 91.7 | 68.5 | |
| 02 | ◎GPT-5.4OpenAI · xHigh | 93.0 | Not reported | Not reported | 75.1 | |
| 03 | ✳Claude Opus 4.6Anthropic · Max | 91.3 | 80.8 | 88.8 | 65.4 | |
| 04 | KKimi K2.6Moonshot AI · ThinkingOpen weights | 90.5 | 80.2 | 89.6 | 66.7 | |
| 05 | DDeepSeek V4 FlashDeepSeek · MaxOpen weights | 88.1 | 79.0 | 91.6 | 56.9 | |
| 06 | ZGLM-5.1Z.ai · ThinkingOpen weights | 86.2 | Not reported | Not reported | 63.5 |
Percentage of real GitHub issues resolved. Results depend on the agent framework, tools, and sampling setup.
Graduate-level science questions, measured with Pass@1.
February 2026 HMMT math problems, measured with Pass@1. Do not combine with AIME or other years.
Multidisciplinary knowledge and reasoning using exact match (EM), distinct from the original MMLU.
Code generation using source-reported Pass@1. Missing public results are marked as not reported.
Accuracy on multi-step terminal tasks. Results depend on the agent setup and tool permissions.
Software engineering tasks resolved, using a different task set from SWE-bench Verified.
This table uses HLE Pass@1 without tools. Do not mix it with tool-enabled results.
Complex questions requiring browsing and multi-step research, measured with Pass@1.
Multi-round retrieval over 1M-token context (MMR). Missing values mean not reported, not zero.
MODEL NOTEBOOK
8 source snapshots from DS-V4-HUUB. Unreviewed specifications are not used as current benchmark evidence.
DeepSeek
Source snapshot · Review required →Anthropic
Source snapshot · Review required →OpenAI
Source snapshot · Review required →MiniMax
Source snapshot · Review required →xAI
Source snapshot · Review required →Alibaba
Source snapshot · Review required →Zhipu AI
Source snapshot · Review required →