All knowledge base guides

EVALUATE & IMPROVE / FIELD GUIDE

Know whether a RAG change actually helped

Use paired questions, source checks, and measured usage to compare two configurations.

7 min read · SaveMyToken editorial · Documentation reviewed 2026-10-05 · Build in your own stack

STEP 01

Write questions before tuning

Build a small test set from actual tasks. For each question, note the required facts and source passages. Keep some questions aside until after choosing a configuration.

  1. Include a direct lookup, paraphrase, identifier, cross-section question, and question with no answer in the collection.
  2. Define what counts as an acceptable answer. An uncovered question passes only if it handles missing evidence appropriately.
  3. Label tuning and held-out validation questions. Do not repeatedly optimize on the held-out set.

STEP 02

Change one thing and keep the rest fixed

Record document versions, model IDs, prompt, retrieval settings, and the changed parameter. Start a fresh chat for comparable single-turn tests.

  1. Capture the baseline's retrieved evidence, answer, and citations before applying the candidate configuration.
  2. Run the same questions against the candidate. Inspect whether the answer-bearing passages reach the model.
  3. Review answers yourself against the expected facts. A retrieval similarity score or an automatic judge score is not an accuracy measurement.

STEP 03

Record performance with its scope

Use actual request traces or provider usage. A blank field means unmeasured, not zero. The comparison worksheet only calculates differences where both configurations have a measurement for the same question.

  1. Record end-to-end milliseconds and actual input/output tokens. Repeat runs separately when latency is variable.
  2. For cost, use the same currency and billing scope for both configurations. Include reranking and query embedding consistently if you count them.
  3. Track one-time indexing, storage, and subscriptions separately from per-query measurements. The worksheet does not estimate these charges.

STEP 04

Validate the selected configuration

Use the questions you held back and inspect failures, including cases where the source cannot answer. The goal is a useful, supported response within your latency and cost budget.

  1. Review paired outcomes instead of treating an aggregate score as proof.
  2. Save chosen retrieval settings in your search service and answer builder, then repeat the full request.
  3. Re-test the full app after applying settings. Repeat this process when documents, models, or prompts change.
Give your changes a fair test.

Use the same questions, inspect the evidence, and record what changed.

Open comparison worksheet