← Evaluation & quality

FREE AI FDE PRACTICE / EVALUATION & QUALITY

An LLM judge rewards longer answers. Can you trust it?

Advanced 10-min practiceEditorial review: 2026-10-05

YOUR SCENARIO

How would you approach this?

Your automatic evaluator prefers a verbose answer that adds an unsupported claim over a concise correct answer. Its scores disagree with support reviewers. How would you repair the evaluation?

This is an illustrative practice scenario. State any additional assumptions in your answer.

Make your case first.

Clarify the goal, identify the biggest uncertainty, outline an approach, and explain how you would test it. Spend about 10 minutes before opening the reference.

Your notes are not submitted or saved. Keep a copy before leaving this page.

Reveal reference approach Clarifying questions, decisions, and tradeoffs

Clarify before designing.

  • What dimensions is the judge supposed to measure, and does it have the actual source evidence?
  • How were its scores compared with human labels across different error types?

One defensible approach

  1. 01

    Separate the criteria

    Create a rubric for support, completeness, relevance, and style instead of one vague quality score. Add concrete positive and negative examples, including fluent answers with false claims. Decide how critical errors affect acceptance.

  2. 02

    Measure judge behavior

    Use an independently reviewed set. For pairwise grading, swap answer order and check consistency. Examine false acceptances, false rejections, and disagreements by category. Mask irrelevant model identity and test whether length alone changes a decision.

  3. 03

    Version and audit the grader

    Record the judge model, prompt, rubric, and evidence version. Recalibrate after changes and route important disagreements to people. Treat a model-generated score as a measurement needing validation rather than ground truth.

Explain the tradeoff

Human review is costly and can disagree too. Spend it on rubric design, representative calibration, and consequential disagreements rather than every easy case.

Common mistakes

  • Choosing the judge because it agrees with the preferred model.
  • Using the same examples to tune and claim final evaluator accuracy.

KEEP THE CONVERSATION GOING

Try the follow-ups.

  1. What would you do if human reviewers disagree?
  2. Which checks could be deterministic instead?

Review your own answer.

Tick the points you covered. This is a reflection checklist, not an automated score or a hiring prediction.

Check the underlying concepts.

The scenario and reference approach were written for SaveMyToken. These sources support the technical concepts; they do not report this question being asked by an employer.

Anthropic: Demystifying evals for AI agents ↗