01 / Start with the complete result, not the rank
Suppose a chart says ‘Model A: 80% on coding’. You still need the benchmark name and version, task set, metric, number of attempts, model snapshot, and test settings. ‘80% pass@1 on a dated LiveCodeBench slice’ and ‘80% of SWE-bench Verified issues resolved by an agent’ answer different questions. The percentage sign does not make them comparable.
Read any benchmark result as a sentence: On [dataset and split], [exact model and setup] achieved [metric] over [number of tasks or votes] under [time, tool, and compute limits]. If the source omits a field, mark it unknown instead of filling it in from the model name.
Result = task set + model snapshot + evaluation setup + metric + sample size + date02 / Know what the familiar benchmark names test
Choose a benchmark that resembles the work you care about. Academic questions, contest programs, repository repairs, and human preferences test different behaviors. Even within a family, the named subset matters: GPQA Diamond is a subset of GPQA; SWE-bench Verified is a reviewed subset of SWE-bench. AIME results need the exam year and paper, while LiveCodeBench results need the problem date range or dataset version.
MMLU-Pro ↗GPQA ↗LiveCodeBench ↗SWE-bench Verified ↗
| Benchmark | What a result mainly describes | Check before comparing |
|---|---|---|
| MMLU-Pro | Accuracy on challenging, ten-choice questions across academic subjects | MMLU versus MMLU-Pro; prompt and reasoning format |
| GPQA Diamond | Accuracy on a difficult expert-written science-question subset | Diamond versus other GPQA subsets; tool or search access |
| AIME | Correct answers on contest mathematics problems | Year, paper, answer extraction, and number of attempts |
| HumanEval | Code completion accepted by functional tests | pass@k, generated candidates, and test harness |
| LiveCodeBench | Code that passes tests on competition-style problems | Problem date window, difficulty slice, pass@k, and execution limits |
| SWE-bench Verified | Repository issues resolved under a coding-agent setup | Subset, agent harness, retries, tools, and submission rules |
| Arena | Human preference in pairwise response comparisons | Arena category, vote count, rating method, and uncertainty |
03 / Translate the score label before comparing numbers
Accuracy is the share of questions graded correct under that benchmark's answer parser. Exact match accepts only an answer that matches the expected form after the evaluator's normalization; a semantically good answer can fail a strict parser. Coding tests instead run generated programs or patches. The metric tells you what counted as success, not how often users will be satisfied.
Pass@1 is success with one sampled candidate. Pass@k asks whether at least one of k candidates succeeds; it can benefit from more attempts and does not describe a single response. A self-consistency or majority-vote result uses yet another selection rule. On SWE-bench, ‘% resolved’ counts issues that pass the prescribed tests after the submitted patch is evaluated. On a preference arena, the displayed rating summarizes pairwise votes; it is neither percent correct nor a probability that your next answer will be right.
HumanEval pass@k implementation ↗SWE-bench scoring ↗Arena rating method ↗
| Label you may see | Read it as | Common trap |
|---|---|---|
| Accuracy / exact match | Fraction of answers accepted by a specified grader | Different parsers or subsets change the denominator |
| pass@1 | One generated candidate passes the check | A sampled estimate may use repeated runs |
| pass@k | At least one of k candidates passes | Extra attempts cost time and money |
| % resolved | Share of repository issues accepted by tests | Agent setup and retries contribute to the result |
| Arena rating | Relative preference estimated from paired votes | A rating is not an accuracy percentage |
04 / Check the denominator and the size of the gap
A two-point lead on 500 questions is only ten additional accepted answers. It may be meaningful, but the headline alone cannot show whether it persists with different questions, prompts, or random samples. Look for the task count, repeated-run results, confidence intervals, and performance by subject or difficulty. A 15-question AIME paper is especially sensitive to one answer: one more correct answer changes the score by about 6.7 percentage points.
When a leaderboard shows adjacent ranks with overlapping uncertainty, treat the order as a weak signal. Do not call models tied solely because two separate confidence intervals overlap; a direct estimate of their difference is more informative. Also check whether failed requests, timeouts, and invalid outputs were counted in the denominator.
05 / Hold the evaluation setup constant
A fair comparison uses the same dataset slice and grading code, then records what each model was allowed to do. Check the exact model or API snapshot; zero-shot versus few-shot examples; system prompt; temperature and sampling; reasoning or thinking budget; maximum output length; context actually supplied; search, browser, code execution, and other tools; and retry or selection policy. A nominal context-window size does not tell you how much context the evaluation supplied or whether the model used it well.
For agent benchmarks, also record the scaffold, test visibility, repository environment, timeout, and token or dollar budget. A coding-agent score measures the model working inside that setup. Comparing two agent runs with different tools or retry counts does not isolate model quality.
SWE-bench submission checklist ↗LiveCodeBench evaluation options ↗
06 / Read speed and price beside quality
Time to first token (TTFT) measures the wait until output begins; time to first answer token can be longer for a reasoning model that emits thinking tokens first. Output tokens per second measures generation after the first token. Neither number alone is total time to a usable answer. End-to-end latency includes the entire request, and agent tasks add tool calls and retries.
Check input length, requested output length, concurrency, serving provider, region, and whether the reported speed is a median or a tail percentile. Price per million tokens is a tariff, not cost per successful task. For an actual workflow, include input, cached input, output and any billed reasoning tokens, tool costs, retries, and the share of tasks that pass your acceptance check. Different tokenizers can bill different token counts for the same text.
07 / Ask whether the test still measures fresh skill
Public tasks can appear in training material, examples, or evaluation practice. A strong score on a familiar set may still describe performance on that set, but it is weaker evidence of transfer to new tasks. Check when tasks were released relative to the model, whether the benchmark has held-out or recently collected problems, and whether answer leakage or flawed tests have been audited. Newer is not automatically cleaner; read the dataset's own documentation.
The benchmark itself can age. OpenAI's later review of SWE-bench Verified reports contamination and test-design limits for evaluating frontier coding models. That does not erase every historical result; it changes how much weight a current selection decision should place on that benchmark alone.
LiveCodeBench date-filtered evaluation ↗Review of SWE-bench Verified limitations ↗
08 / Work through a claim before accepting it
Consider the fictional claim ‘Model B beats Model A at coding: 90% versus 84%.’ If B's number is pass@8 on one LiveCodeBench date range while A's is pass@1 on another, the comparison has two confounders before model quality enters the picture: more attempts and different problems. Ask for pass@1 on the same date range and difficulty slice, with the same language, time limit, prompts, and execution rules. Then compare uncertainty, latency, and cost per accepted solution.
Use public benchmarks to decide which models deserve a trial. For the final choice, run a small set of your own representative tasks with one written success rule and the same budget for each model. Keep missing public scores as unknown, and resist folding unlike metrics into a single unexplained ‘intelligence’ number.
Before trusting a row: What tasks? Which version? What metric? How many tries? Which tools and budget? How many samples? When was it run? What did a successful result cost?