LANGUAGE & CODING / DeepSeek
DeepSeek V4.1 Flash (fast)
Multimodal reasoning and tool workflows with peak and off-peak API rates.
PRICE / USD
$0.3
Per 1 million output tokens · input billed separately
- Input / 1M tokens
- $0.3
- Output / 1M tokens
- $1.2
- Cached input / 1M tokens
- $0.006
- Cache write / 1M tokens
- Not verified
Off-peak rates are 50% lower. Peak: weekdays 01:00–04:00 and 06:00–10:00 UTC.
DeepSeek models & pricing ↗SPEED & CAPABILITIES
214.1 tokens/s
DeepSeek API · V4.1 Flash · Max reasoning. Artificial Analysis reports a rolling 72-hour median of output speed after the first token, not total response time. Checked 2026-10-05; speed varies with workload and serving conditions.
Artificial Analysis · DeepSeek V4.1 Flash (Max) ↗Artificial Analysis · performance methodology v2.2.0 ↗ · Checked 2026-10-05. (fast) means median output above 150 tokens/s in these settings.
- Model score
- 39
- Our own output-speed measurement
- Not measured
- Context window
- 1,000,000
- API model ID
deepseek-flash
Artificial Analysis tested DeepSeek V4.1 Flash · max. This mostly English, text-based index is not a guarantee of performance on your workload. Artificial Analysis · DeepSeek V4.1 Flash · max ↗
The old V4 Flash API aliases now serve V4.1 Flash. Preview benchmark scores are not transferred to this version.
Suggested starting role: everyday coding. A low-cost starting point for code changes and tool workflows, based on its documented reasoning and agent support. Check your own tests before using it for larger changes. This is editorial guidance, not a benchmark score.
Sources & methodology
DeepSeek models & pricing ↗ Checked 2026-10-05
Artificial Analysis · DeepSeek V4.1 Flash (Max) ↗ Checked 2026-10-05
Artificial Analysis · performance methodology v2.2.0 ↗ Checked 2026-10-05
Artificial Analysis · Intelligence Index methodology v4.3.2 ↗ Checked 2026-10-05