WHAT YOU WILL BUILD
A three-row evaluation showing the tradeoff between evidence coverage and extra context.
The example question needs both a rule and an exception. Precision is relevant retrieved items divided by retrieved items; recall is relevant retrieved items divided by all labeled relevant items. They describe retrieval, not the truthfulness of a final answer.
Before you start
Use a text editor and Node.js 22 or newer ↗. Check your version with node --version. Save the downloaded file in an empty folder, then open a terminal in that folder.
All data is included. No packages, account, or API key are required. The labels and rankings are a teaching fixture. The results are arithmetic examples, not a model benchmark or a production measurement.
Download the lab (.mjs)FOLLOW ALONG
Work through the example.
- 01
Fix the relevance labels
Mark rule and exception as relevant; shipping is irrelevant to this question. Do this before changing K, so the target cannot move with your result.
- 02
Run K from one to three
At K=1 precision is 1.00 but recall is only 0.50. At K=3 recall reaches 1.00 while precision drops to 0.67 because shipping is included.
- 03
Check the denominator
For K=2, only one of the two retrieved items is relevant, so precision is 1/2. There are two relevant items overall, so recall is also 1/2.
- 04
Build a small comparison sheet
Add several representative questions with their relevant document IDs. Compare two configurations on the same questions. Record per-question failures before averaging, then review the answers and measure latency separately.
THE COMPLETE LAB
Run it locally.
Run this command from the folder containing your downloaded file:
node rag-evaluation.mjsView or copy the complete JavaScript
// SaveMyToken local lab. Run with Node.js 22 or newer.
import assert from "node:assert/strict";
const relevant = new Set(["rule", "exception"]);
const ranked = ["rule", "shipping", "exception"];
function evaluate(k) {
const retrieved = [...new Set(ranked)].slice(0, k);
const hits = retrieved.filter(id => relevant.has(id)).length;
return { precision: hits / retrieved.length, recall: hits / relevant.size };
}
for (const k of [1, 2, 3]) {
const result = evaluate(k);
console.log("K=" + k + " precision=" + result.precision.toFixed(2) + " recall=" + result.recall.toFixed(2));
}
assert.equal(evaluate(1).recall, 0.5);
assert.equal(evaluate(3).recall, 1);
assert.equal(evaluate(3).precision, 2 / 3);
Expected output for the unchanged example
K=1 precision=1.00 recall=0.50
K=2 precision=0.50 recall=0.50
K=3 precision=0.67 recall=1.00The assertions also check the baseline behavior. After changing an input, predict the result and update the relevant assertion.
NOW CHANGE ONE THING
Make the example your own.
Move exception to second place. Predict precision and recall at K=2, then update the fixture and assertions. Explain why better ordering can reduce the context needed.
You are done when…
You can compute all three rows by hand and identify a case where perfect precision still misses required evidence.
If something goes wrong
For a query with no relevant documents, recall has a zero denominator. Evaluate correct abstention separately instead of dividing by zero. This lab uses only nonempty labels and result sets.
EXPLAIN WHAT YOU LEARNED
Interview practice
Try answering aloud before opening the reference answer. These are original learning questions, not a record of any employer's interviews.
01Can precision be perfect while retrieval is incomplete?
Yes. Retrieving only one of two required passages gives precision 1 and recall 0.5. A missing exception may still make the final answer wrong.
Watch for: A high precision number does not establish answer completeness.
02Does increasing K always help?
It can improve coverage but add irrelevant text and processing cost. Compare evidence coverage, final answers, and latency under the same conditions.
Watch for: More context can include outdated or conflicting passages.
03How should you compare two RAG configurations?
Use the same source snapshot, questions, labels, and evaluation method. Change one setting when diagnosing a cause, and keep a separate held-out set for the final comparison.
Watch for: Repeated tuning on the final test set inflates confidence in the result.
04Which costs belong in a RAG evaluation?
Track ingestion and indexing separately from per-query retrieval, reranking, and generation. Record retries and cache behavior rather than counting only the final model response.
Watch for: A local arithmetic example is not evidence of real API savings.
05How do you handle questions with no answer in the corpus?
Include them deliberately and measure appropriate abstention and false claims. Keep that result separate from recall over queries that have labeled relevant evidence.
Watch for: Dropping uncovered questions hides a common production failure.
Sources and further reading
The explanation and local exercises were written for SaveMyToken. These references support the underlying concepts; the sample outputs describe only the supplied examples.
Hugging Face: evaluating RAG ↗