← Choose models

Language & coding

What mattersDeepSeek V4.1 Flash (fast) ↗DeepSeek
Use casesCoding · Agents & tools · Reasoning · Visual understanding
Model score39DeepSeek V4.1 Flash · maxArtificial Analysis · DeepSeek V4.1 Flash · max ↗
Output / 1M tokens · USD$0.3
Published rates · USD$0.3 input / $1.2 outputPer 1M native tokens
Cached input / 1M$0.006
Cache write / 1MNot verified
Pricing conditionsDeepSeek API · peak hoursOff-peak rates are 50% lower. Peak: weekdays 01:00–04:00 and 06:00–10:00 UTC.DeepSeek models & pricing ↗
Output speed · tokens/s214.1 tokens/sDeepSeek API · V4.1 Flash · Max reasoning. Artificial Analysis reports a rolling 72-hour median of output speed after the first token, not total response time. Checked 2026-10-05; speed varies with workload and serving conditions.Artificial Analysis · DeepSeek V4.1 Flash (Max) ↗Artificial Analysis · performance methodology v2.2.0 ↗ · Checked 2026-10-05
Context window1,000,000
Model IDdeepseek-flash
Version notesThe old V4 Flash API aliases now serve V4.1 Flash. Preview benchmark scores are not transferred to this version.
Official documentationDeepSeek models & pricing ↗Checked 2026-10-05
Language prices are USD per million tokens; input and output are billed separately. Video samples use 5s of 768P output where verified. Extra tools, reasoning, cache writes, input media, storage, and taxes can change the total. (fast) marks a published median output speed above 150 tokens/s in the linked test settings.