01 / Start with the complete result, not the rank

Suppose a chart says ‘Model A: 80% on coding’. You still need the benchmark name and version, task set, metric, number of attempts, model snapshot, and test settings. ‘80% pass@1 on a dated LiveCodeBench slice’ and ‘80% of SWE-bench Verified issues resolved by an agent’ answer different questions. The percentage sign does not make them comparable.

Read any benchmark result as a sentence: On [dataset and split], [exact model and setup] achieved [metric] over [number of tasks or votes] under [time, tool, and compute limits]. If the source omits a field, mark it unknown instead of filling it in from the model name.

Result = task set + model snapshot + evaluation setup + metric + sample size + date

02 / Know what the familiar benchmark names test

Choose a benchmark that resembles the work you care about. Academic questions, contest programs, repository repairs, and human preferences test different behaviors. Even within a family, the named subset matters: GPQA Diamond is a subset of GPQA; SWE-bench Verified is a reviewed subset of SWE-bench. AIME results need the exam year and paper, while LiveCodeBench results need the problem date range or dataset version.

MMLU-Pro ↗GPQA ↗LiveCodeBench ↗SWE-bench Verified ↗

Know what the familiar benchmark names test
BenchmarkWhat a result mainly describesCheck before comparing
MMLU-ProAccuracy on challenging, ten-choice questions across academic subjectsMMLU versus MMLU-Pro; prompt and reasoning format
GPQA DiamondAccuracy on a difficult expert-written science-question subsetDiamond versus other GPQA subsets; tool or search access
AIMECorrect answers on contest mathematics problemsYear, paper, answer extraction, and number of attempts
HumanEvalCode completion accepted by functional testspass@k, generated candidates, and test harness
LiveCodeBenchCode that passes tests on competition-style problemsProblem date window, difficulty slice, pass@k, and execution limits
SWE-bench VerifiedRepository issues resolved under a coding-agent setupSubset, agent harness, retries, tools, and submission rules
ArenaHuman preference in pairwise response comparisonsArena category, vote count, rating method, and uncertainty

03 / Translate the score label before comparing numbers

Accuracy is the share of questions graded correct under that benchmark's answer parser. Exact match accepts only an answer that matches the expected form after the evaluator's normalization; a semantically good answer can fail a strict parser. Coding tests instead run generated programs or patches. The metric tells you what counted as success, not how often users will be satisfied.

Pass@1 is success with one sampled candidate. Pass@k asks whether at least one of k candidates succeeds; it can benefit from more attempts and does not describe a single response. A self-consistency or majority-vote result uses yet another selection rule. On SWE-bench, ‘% resolved’ counts issues that pass the prescribed tests after the submitted patch is evaluated. On a preference arena, the displayed rating summarizes pairwise votes; it is neither percent correct nor a probability that your next answer will be right.

HumanEval pass@k implementation ↗SWE-bench scoring ↗Arena rating method ↗

Translate the score label before comparing numbers
Label you may seeRead it asCommon trap
Accuracy / exact matchFraction of answers accepted by a specified graderDifferent parsers or subsets change the denominator
pass@1One generated candidate passes the checkA sampled estimate may use repeated runs
pass@kAt least one of k candidates passesExtra attempts cost time and money
% resolvedShare of repository issues accepted by testsAgent setup and retries contribute to the result
Arena ratingRelative preference estimated from paired votesA rating is not an accuracy percentage

04 / Check the denominator and the size of the gap

A two-point lead on 500 questions is only ten additional accepted answers. It may be meaningful, but the headline alone cannot show whether it persists with different questions, prompts, or random samples. Look for the task count, repeated-run results, confidence intervals, and performance by subject or difficulty. A 15-question AIME paper is especially sensitive to one answer: one more correct answer changes the score by about 6.7 percentage points.

When a leaderboard shows adjacent ranks with overlapping uncertainty, treat the order as a weak signal. Do not call models tied solely because two separate confidence intervals overlap; a direct estimate of their difference is more informative. Also check whether failed requests, timeouts, and invalid outputs were counted in the denominator.

MAA AIME format ↗Arena rating and confidence intervals ↗

05 / Hold the evaluation setup constant

A fair comparison uses the same dataset slice and grading code, then records what each model was allowed to do. Check the exact model or API snapshot; zero-shot versus few-shot examples; system prompt; temperature and sampling; reasoning or thinking budget; maximum output length; context actually supplied; search, browser, code execution, and other tools; and retry or selection policy. A nominal context-window size does not tell you how much context the evaluation supplied or whether the model used it well.

For agent benchmarks, also record the scaffold, test visibility, repository environment, timeout, and token or dollar budget. A coding-agent score measures the model working inside that setup. Comparing two agent runs with different tools or retry counts does not isolate model quality.

SWE-bench submission checklist ↗LiveCodeBench evaluation options ↗

06 / Read speed and price beside quality

Time to first token (TTFT) measures the wait until output begins; time to first answer token can be longer for a reasoning model that emits thinking tokens first. Output tokens per second measures generation after the first token. Neither number alone is total time to a usable answer. End-to-end latency includes the entire request, and agent tasks add tool calls and retries.

Check input length, requested output length, concurrency, serving provider, region, and whether the reported speed is a median or a tail percentile. Price per million tokens is a tariff, not cost per successful task. For an actual workflow, include input, cached input, output and any billed reasoning tokens, tool costs, retries, and the share of tasks that pass your acceptance check. Different tokenizers can bill different token counts for the same text.

API performance definitions and workloads ↗

07 / Ask whether the test still measures fresh skill

Public tasks can appear in training material, examples, or evaluation practice. A strong score on a familiar set may still describe performance on that set, but it is weaker evidence of transfer to new tasks. Check when tasks were released relative to the model, whether the benchmark has held-out or recently collected problems, and whether answer leakage or flawed tests have been audited. Newer is not automatically cleaner; read the dataset's own documentation.

The benchmark itself can age. OpenAI's later review of SWE-bench Verified reports contamination and test-design limits for evaluating frontier coding models. That does not erase every historical result; it changes how much weight a current selection decision should place on that benchmark alone.

LiveCodeBench date-filtered evaluation ↗Review of SWE-bench Verified limitations ↗

08 / Work through a claim before accepting it

Consider the fictional claim ‘Model B beats Model A at coding: 90% versus 84%.’ If B's number is pass@8 on one LiveCodeBench date range while A's is pass@1 on another, the comparison has two confounders before model quality enters the picture: more attempts and different problems. Ask for pass@1 on the same date range and difficulty slice, with the same language, time limit, prompts, and execution rules. Then compare uncertainty, latency, and cost per accepted solution.

Use public benchmarks to decide which models deserve a trial. For the final choice, run a small set of your own representative tasks with one written success rule and the same budget for each model. Keep missing public scores as unknown, and resist folding unlike metrics into a single unexplained ‘intelligence’ number.

Before trusting a row: What tasks? Which version? What metric? How many tries? Which tools and budget? How many samples? When was it run? What did a successful result cost?