Decision & classification6
Find the right fitSources checked 2026-10-05

Sample: 5 seconds of 768P output where verified. Reference inputs cost extra. Other resolutions and audio settings are not interchangeable.

7 video models
Kuaishou

Kling Video 3.0

Text-to-video and image-to-video with shot control and native audio.

Text to videoImage to video
5s output sampleNot verifiedQuote depends on settings
SpeedNot verifiedNo reviewed numeric benchmark

Official API quote not verified

MiniMax

MiniMax H3

Generate or edit video using text, frames, and audiovisual references.

Text to videoImage to videoReference to video
5s output · 768P$0.4$0.08 / output second
SpeedNot verifiedNo reviewed numeric benchmark

MiniMax API · output only

MiniMax / fal.ai

MiniMax H3 Max

A speed-focused H3 variant for text, image, and reference-based video.

Text to videoImage to videoReference to video
5s output · 768P$0.4$0.08 / output second
SpeedFaster H3 variantProvider-reported

MiniMax API · output only

ByteDance

Seedance 2.0

The quality-focused 2.0 variant for multimodal video creation and editing.

Text to videoImage to videoReference to video
5s output sampleToken-priced$7 / 1M video tokens · no video input
SpeedNot verifiedNo reviewed numeric benchmark

BytePlus ModelArk · list rates / 1M video tokens

ByteDance

Seedance 2.0 Fast

The faster 2.0 variant for text, image, reference, and video-editing workflows.

Text to videoImage to videoReference to video
5s output sampleToken-priced$5.6 / 1M video tokens · no video input
SpeedFaster 2.0 variantProvider-reported

BytePlus ModelArk · list rates / 1M video tokens

ByteDance

Seedance 2.0 Mini

The lower-cost 2.0 variant for video drafts, references, and editing.

Text to videoImage to videoReference to video
5s output sampleToken-priced$3.5 / 1M video tokens · no video input
SpeedNot verifiedNo reviewed numeric benchmark

BytePlus ModelArk · list rates / 1M video tokens

ByteDance

Seedance 2.5

Audio-video generation, multimodal references, and targeted editing.

Text to videoImage to videoReference to video
5s output sampleNot verifiedQuote depends on settings
SpeedNot verifiedNo reviewed numeric benchmark

BytePlus ModelArk · token-based video billing

Use-case tags are editorial groupings based on official capabilities. Language models earn (fast) only with a published median output speed above 150 tokens/s in the linked test settings. Video speed labels remain provider-reported. Unknown values are excluded from price bands and sorted last.

Model evaluation & benchmark research

Original research and practical experience, with short notes to help you choose.

Model benchmarks

Benchmarking GPT-6 Astra

A dated evaluation of GPT-6 Astra across capability, reasoning settings, token use, and task cost.

Artificial Analysis · Model evaluation
Model benchmarks

Factuality in the Arena

Why user preference alone does not settle factual accuracy, and how Arena adds factuality signals.

Arena · Evaluation methodology
Historical benchmarks · 6 models / 10 metrics
DeepSeek V4 report · Historical versions and reasoning modes
#Model / reasoning modeCompare
01✦Gemini 3.1 ProGoogle · High94.3
80.691.768.5
02◎GPT-5.4OpenAI · xHigh93.0
Not reportedNot reported75.1
03✳Claude Opus 4.6Anthropic · Max91.3
80.888.865.4
04KKimi K2.6Moonshot AI · ThinkingOpen weights90.5
80.289.666.7
05DDeepSeek V4 FlashDeepSeek · MaxOpen weights88.1
79.091.656.9
06ZGLM-5.1Z.ai · ThinkingOpen weights86.2
Not reportedNot reported63.5
GPQA Diamond · Pass@1 % · Provider-reported data, not our own tests. Reasoning budgets may differ.Source data
Historical snapshot checked on September 24, 2026, from the DeepSeek V4 preview model card. These are provider-reported results, not our own tests or necessarily the latest model versions. Reasoning budgets and agent frameworks may differ. Unreported metrics remain unknown.

What does each metric measure?

SWE-bench Verified / Resolved %

Percentage of real GitHub issues resolved. Results depend on the agent framework, tools, and sampling setup.

GPQA Diamond / Pass@1 %

Graduate-level science questions, measured with Pass@1.

HMMT 2026 Feb / Pass@1 %

February 2026 HMMT math problems, measured with Pass@1. Do not combine with AIME or other years.

MMLU-Pro / EM %

Multidisciplinary knowledge and reasoning using exact match (EM), distinct from the original MMLU.

LiveCodeBench / Pass@1 %

Code generation using source-reported Pass@1. Missing public results are marked as not reported.

Terminal-Bench 2.0 / Accuracy %

Accuracy on multi-step terminal tasks. Results depend on the agent setup and tool permissions.

SWE-bench Pro / Resolved %

Software engineering tasks resolved, using a different task set from SWE-bench Verified.

Humanity’s Last Exam / Pass@1 %

This table uses HLE Pass@1 without tools. Do not mix it with tool-enabled results.

BrowseComp / Pass@1 %

Complex questions requiring browsing and multi-step research, measured with Pass@1.

MRCR 1M / MMR %

Multi-round retrieval over 1M-token context (MMR). Missing values mean not reported, not zero.

Imported model notebook · unreviewed historical material

MODEL NOTEBOOK

Historical model notes

8 source snapshots from DS-V4-HUUB. Unreviewed specifications are not used as current benchmark evidence.