Decision & classification6
Find the right fitSources checked 2026-10-05

Language model comparisons use the uncached input price in USD per 1 million tokens. Output is billed separately and shown on each card. Excludes tools, cache writes, storage, and tax.

TASK LEVEL × INPUT PRICE

Compare model roles and prices

Each model keeps its exact score height. The right side groups those scores into complex planning, everyday coding, and simple tasks.

Editorial score guide
Uncached input · USD per 1,000,000 tokens · output priced separately
Lower cost ≤ $0.5 Mid-range ≤ $2 Higher budget > $2(fast) = median output > 150 tokens/s in the tested settingsSwipe chart to explore all models →
Language model score and task tier versus input priceHorizontal position is USD per one million input tokens on a logarithmic scale. Vertical position is the exact model score. The right-side regions classify scores above 50 as complex planning, 39 through 50 as everyday coding, and below 39 as simple tasks. MiMo V2.6 Pro uses 38. Unscored models are omitted. A fast suffix means published median output above 150 tokens per second in the tested settings.SCORE ↑← LESS EXPENSIVEMORE EXPENSIVE →TASK TIERComplex planningScore > 50Everyday codingScore 39–50Simple tasksScore < 392030395060$0.05$0.1$0.3$1$3$10$20INPUT PRICE · USD / 1M TOKENS · LOG SCALEClaude Opus 5.5 · score 58 · Complex planning · $4 / 1M input tokens · Not benchmarked · Select for evidenceOpus 5.5Claude Sonnet 5.5 · score 56 · Complex planning · $2 / 1M input tokens · Not benchmarked · Select for evidenceSonnet 5.5GPT-6 Astra · score 53 · Complex planning · $10 / 1M input tokens · Not benchmarked · Select for evidenceGPT-6 AstraGrok 4.7 · score 46 · Everyday coding · $2 / 1M input tokens · Not benchmarked · Select for evidenceGrok 4.7GLM-5.3 · score 45 · Everyday coding · $1.4 / 1M input tokens · Not benchmarked · Select for evidenceGLM-5.3Kimi K3 · score 44 · Everyday coding · $3 / 1M input tokens · Not benchmarked · Select for evidenceKimi K3Gemini 3.8 Flash (fast) · score 41 · Everyday coding · $0.75 / 1M input tokens · 236.5 tokens/s · Select for evidenceGemini Flash (fast)DeepSeek V4.1 Flash (fast) · score 39 · Everyday coding · $0.3 / 1M input tokens · 214.1 tokens/s · Select for evidenceDS Flash (fast)GPT-6 Luna · score 38 · Simple tasks · $0.1 / 1M input tokens · Not benchmarked · Select for evidenceGPT-6 LunaMiMo V2.6 Pro · score 38 · Simple tasks · $0.435 / 1M input tokens · 42.8 tokens/s · Select for evidenceMiMo ProMiniMax M3 · score 29 · Simple tasks · $0.3 / 1M input tokens · Not benchmarked · Select for evidenceMiniMax M3

Dot height is the exact score; the right-side regions show its task tier. Labels move downward only to avoid overlap. Select a point for details.

Not plotted because a verified index score or a positive input price is unavailable: MiniMax M2.7 Highspeed, Qwen3.8-Max.

How to read this map

The vertical axis groups models into three score-based rows: above 50 for complex planning, 39–50 for everyday coding, and below 39 for simple tasks. MiMo V2.6 Pro is set conservatively to 38. Scores do not replace testing on your own workload. Task buttons highlight one row without changing model placement.

“Lower cost” describes token prices, not proven value per successful task. For real value, compare correctness, retries, tools, and total billed tokens on the same tasks.

The horizontal axis uses a logarithmic scale: equal spacing means equal price ratios. Prices are USD per one million native tokens, using uncached input under each entry’s stated API conditions. Price filters, sorting, and chart positions use the same input rate. Capability evidence and pricing sources are linked in each detail page.

Model detail pages show a (fast) badge only when a published median output speed exceeds 150 tokens/s in the tested settings. That speed does not include time to first token or total task time. Speed methodology ↗

13 language models
Anthropic

Claude Opus 5.5

Long-running coding agents and knowledge work.

Intelligence Index v4.3.2: 58 · Claude Opus 5.5 · max, default fallback

CodingAgents & toolsReasoning
Input / 1M tokens$4$4 input / $20 output per 1M tokens
SpeedNot benchmarkedNo reviewed numeric benchmark

Claude API · standard processing

Anthropic

Claude Sonnet 5.5

General coding, visual tasks, and document workflows.

Intelligence Index v4.3.2: 56 · Claude Sonnet 5.5 · max, default fallback

CodingAgents & toolsDocuments & writing
Input / 1M tokens$2$2 input / $10 output per 1M tokens
SpeedNot benchmarkedNo reviewed numeric benchmark

Claude API · standard processing

DeepSeek

DeepSeek V4.1 Flash (fast)

Multimodal reasoning and tool workflows with peak and off-peak API rates.

Intelligence Index v4.3.2: 39 · DeepSeek V4.1 Flash · max

CodingAgents & toolsReasoning
Input / 1M tokens$0.3$0.3 input / $1.2 output per 1M tokens
Speed214.1 tokens/sArtificial Analysis · median

DeepSeek API · peak hours

Google

Gemini 3.8 Flash (fast)

Software engineering, autonomous agents, and enterprise workflows.

Intelligence Index v4.3.2: 41 · Gemini 3.8 Flash · high

CodingAgents & toolsReasoning
Input / 1M tokens$0.75$0.75 input / $3.75 output per 1M tokens
Speed236.5 tokens/sArtificial Analysis · median

Gemini Developer API · Standard · introductory

Z.ai

GLM-5.3

Text-only reasoning, complex coding, and long-running agents.

Intelligence Index v4.3.2: 45 · GLM-5.3 · max

CodingAgents & toolsReasoning
Input / 1M tokens$1.4$1.4 input / $4.4 output per 1M tokens
SpeedNot benchmarkedNo reviewed numeric benchmark

Z.ai API · listed USD rates

OpenAI

GPT-6 Astra

Complex reasoning, software engineering, research, and document creation.

Intelligence Index v4.3.2: 53 · GPT-6 Astra · max

CodingAgents & toolsReasoning
Input / 1M tokens$10$10 input / $50 output per 1M tokens
SpeedNot benchmarkedNo reviewed numeric benchmark

OpenAI API · Standard · short context

OpenAI

GPT-6 Luna

Focused, high-volume tasks with low token prices.

Intelligence Index v4.3.2: 38 · GPT-6 Luna · max

Everyday tasksDocuments & writingVisual understanding
Input / 1M tokens$0.1$0.1 input / $0.5 output per 1M tokens
SpeedNot benchmarkedNo reviewed numeric benchmark

OpenAI API · Standard · short context

xAI / SpaceXAI

Grok 4.7

Coding and configurable reasoning with agent tool support.

Intelligence Index v4.3.2: 46 · Grok 4.7 · xhigh

CodingAgents & toolsReasoning
Input / 1M tokens$2$2 input / $6 output per 1M tokens
SpeedNot benchmarkedNo reviewed numeric benchmark

Grok API · listed token rates

Moonshot AI

Kimi K3

Long-horizon software engineering, reasoning, and visual knowledge work.

Intelligence Index v4.3.2: 44 · Kimi K3 · max

CodingAgents & toolsReasoning
Input / 1M tokens$3$3 input / $15 output per 1M tokens
SpeedNot benchmarkedNo reviewed numeric benchmark

Kimi API · listed USD rates

Xiaomi

MiMo V2.6 Pro

Lower-cost multimodal reasoning, coding, and agent workflows; verify demanding plans before use.

Model score: 38 · conservative editorial adjustment

ReasoningAgents & toolsCoding
Input / 1M tokens$0.435$0.435 input / $0.87 output per 1M tokens
Speed42.8 tokens/sArtificial Analysis · median

Xiaomi MiMo API · overseas USD · real-time

MiniMax

MiniMax M2.7 Highspeed

A faster serving option for M2.7 coding and agent workflows.

Intelligence Index: no reviewed score for this version

CodingAgents & toolsDocuments & writing
Input / 1M tokens$0.6$0.6 input / $2.4 output per 1M tokens
SpeedNot benchmarkedNo reviewed numeric benchmark

MiniMax API · Highspeed

MiniMax

MiniMax M3

Multimodal coding and agent tasks with a million-token context.

Intelligence Index v4.3.2: 29 · MiniMax M3 · reasoning

CodingAgents & toolsVisual understanding
Input / 1M tokens$0.3$0.3 input / $1.2 output per 1M tokens
SpeedNot benchmarkedNo reviewed numeric benchmark

MiniMax API · Standard · input ≤512K

Alibaba Cloud

Qwen3.8-Max

Coding, office documents, and visual understanding across long inputs.

Intelligence Index: no reviewed score for this version

CodingAgents & toolsDocuments & writing
Input / 1M tokens$2$2 input / $6 output per 1M tokens
SpeedNot benchmarkedNo reviewed numeric benchmark

Model Studio · Singapore / International

Use-case tags are editorial groupings based on official capabilities. Language models earn (fast) only with a published median output speed above 150 tokens/s in the linked test settings. Video speed labels remain provider-reported. Unknown values are excluded from price bands and sorted last.

Model evaluation & benchmark research

Original research and practical experience, with short notes to help you choose.

Model benchmarks

Benchmarking GPT-6 Astra

A dated evaluation of GPT-6 Astra across capability, reasoning settings, token use, and task cost.

Artificial Analysis · Model evaluation
Model benchmarks

Factuality in the Arena

Why user preference alone does not settle factual accuracy, and how Arena adds factuality signals.

Arena · Evaluation methodology
Historical benchmarks · 6 models / 10 metrics
DeepSeek V4 report · Historical versions and reasoning modes
#Model / reasoning modeCompare
01✦Gemini 3.1 ProGoogle · High94.3
80.691.768.5
02◎GPT-5.4OpenAI · xHigh93.0
Not reportedNot reported75.1
03✳Claude Opus 4.6Anthropic · Max91.3
80.888.865.4
04KKimi K2.6Moonshot AI · ThinkingOpen weights90.5
80.289.666.7
05DDeepSeek V4 FlashDeepSeek · MaxOpen weights88.1
79.091.656.9
06ZGLM-5.1Z.ai · ThinkingOpen weights86.2
Not reportedNot reported63.5
GPQA Diamond · Pass@1 % · Provider-reported data, not our own tests. Reasoning budgets may differ.Source data
Historical snapshot checked on September 24, 2026, from the DeepSeek V4 preview model card. These are provider-reported results, not our own tests or necessarily the latest model versions. Reasoning budgets and agent frameworks may differ. Unreported metrics remain unknown.

What does each metric measure?

SWE-bench Verified / Resolved %

Percentage of real GitHub issues resolved. Results depend on the agent framework, tools, and sampling setup.

GPQA Diamond / Pass@1 %

Graduate-level science questions, measured with Pass@1.

HMMT 2026 Feb / Pass@1 %

February 2026 HMMT math problems, measured with Pass@1. Do not combine with AIME or other years.

MMLU-Pro / EM %

Multidisciplinary knowledge and reasoning using exact match (EM), distinct from the original MMLU.

LiveCodeBench / Pass@1 %

Code generation using source-reported Pass@1. Missing public results are marked as not reported.

Terminal-Bench 2.0 / Accuracy %

Accuracy on multi-step terminal tasks. Results depend on the agent setup and tool permissions.

SWE-bench Pro / Resolved %

Software engineering tasks resolved, using a different task set from SWE-bench Verified.

Humanity’s Last Exam / Pass@1 %

This table uses HLE Pass@1 without tools. Do not mix it with tool-enabled results.

BrowseComp / Pass@1 %

Complex questions requiring browsing and multi-step research, measured with Pass@1.

MRCR 1M / MMR %

Multi-round retrieval over 1M-token context (MMR). Missing values mean not reported, not zero.

Imported model notebook · unreviewed historical material

MODEL NOTEBOOK

Historical model notes

8 source snapshots from DS-V4-HUUB. Unreviewed specifications are not used as current benchmark evidence.