How to use these results

This historical compilation from the DeepSeek V4 model card provides a traceable starting point. It is not an independent retest or proof of identical reasoning budgets, tool setups, or costs.

Verify on real tasks

Record completion rate, retries, latency, and total cost on the same tasks. Missing public results are unknown and should not be filled using another model or release.

Original evaluation

DeepSeek V4 model card · Evaluation results ↗

Scores retain the reported version and reasoning mode. They do not automatically change when a similarly named model is updated.