Kling Video 3.0
Text-to-video and image-to-video with shot control and native audio.
Official API quote not verified
COMPARE MODELS
Explore video models by output settings, price conditions, and documented evidence. Check the version and provider source before you choose.
Sample: 5 seconds of 768P output where verified. Reference inputs cost extra. Other resolutions and audio settings are not interchangeable.
Text-to-video and image-to-video with shot control and native audio.
Official API quote not verified
Generate or edit video using text, frames, and audiovisual references.
MiniMax API · output only
A speed-focused H3 variant for text, image, and reference-based video.
MiniMax API · output only
The quality-focused 2.0 variant for multimodal video creation and editing.
BytePlus ModelArk · list rates / 1M video tokens
The faster 2.0 variant for text, image, reference, and video-editing workflows.
BytePlus ModelArk · list rates / 1M video tokens
The lower-cost 2.0 variant for video drafts, references, and editing.
BytePlus ModelArk · list rates / 1M video tokens
Audio-video generation, multimodal references, and targeted editing.
BytePlus ModelArk · token-based video billing
Use-case tags are editorial groupings based on official capabilities. Language models earn (fast) only with a published median output speed above 150 tokens/s in the linked test settings. Video speed labels remain provider-reported. Unknown values are excluded from price bands and sorted last.
Original research and practical experience, with short notes to help you choose.
A comparison of model–harness combinations that puts task success and cost next to each other.
How domain-specific benchmark slices and weights change the meaning of a model capability score.
A dated evaluation of GPT-6 Astra across capability, reasoning settings, token use, and task cost.
Why user preference alone does not settle factual accuracy, and how Arena adds factuality signals.
| # | Model / reasoning mode | Compare | ||||
|---|---|---|---|---|---|---|
| 01 | ✦Gemini 3.1 ProGoogle · High | 94.3 | 80.6 | 91.7 | 68.5 | |
| 02 | ◎GPT-5.4OpenAI · xHigh | 93.0 | Not reported | Not reported | 75.1 | |
| 03 | ✳Claude Opus 4.6Anthropic · Max | 91.3 | 80.8 | 88.8 | 65.4 | |
| 04 | KKimi K2.6Moonshot AI · ThinkingOpen weights | 90.5 | 80.2 | 89.6 | 66.7 | |
| 05 | DDeepSeek V4 FlashDeepSeek · MaxOpen weights | 88.1 | 79.0 | 91.6 | 56.9 | |
| 06 | ZGLM-5.1Z.ai · ThinkingOpen weights | 86.2 | Not reported | Not reported | 63.5 |
Percentage of real GitHub issues resolved. Results depend on the agent framework, tools, and sampling setup.
Graduate-level science questions, measured with Pass@1.
February 2026 HMMT math problems, measured with Pass@1. Do not combine with AIME or other years.
Multidisciplinary knowledge and reasoning using exact match (EM), distinct from the original MMLU.
Code generation using source-reported Pass@1. Missing public results are marked as not reported.
Accuracy on multi-step terminal tasks. Results depend on the agent setup and tool permissions.
Software engineering tasks resolved, using a different task set from SWE-bench Verified.
This table uses HLE Pass@1 without tools. Do not mix it with tool-enabled results.
Complex questions requiring browsing and multi-step research, measured with Pass@1.
Multi-round retrieval over 1M-token context (MMR). Missing values mean not reported, not zero.
MODEL NOTEBOOK
8 source snapshots from DS-V4-HUUB. Unreviewed specifications are not used as current benchmark evidence.
DeepSeek
Source snapshot · Review required →Anthropic
Source snapshot · Review required →OpenAI
Source snapshot · Review required →MiniMax
Source snapshot · Review required →xAI
Source snapshot · Review required →Alibaba
Source snapshot · Review required →Zhipu AI
Source snapshot · Review required →