03 / VERIFY

Evaluate your RAG changes.

Use a simple comparison worksheet to capture the evidence. Run the same questions against two configurations in your own workspace.

Your observations, kept in this page.

Run tests in your own RAG app, then enter the results here. Nothing in this worksheet is uploaded or automatically scored. Download your CSV before leaving; page entries are not saved. Read the testing guide first →

1. Record what changed

Keep documents, models, and prompts fixed unless they are the change you are testing.

Use the same currency and billing scope for both runs. Currency selection labels your entries; it does not convert values. Leave unknown measurements blank.

2. Compare the same questions

Include answerable and unanswerable questions. Reserve some for validation.

1 / 10 questions

Question 1

Baseline run
Candidate run

3. Review paired results

Only questions with a name and measurements on both sides count.

Paired tuning results for baseline and candidate
MeasureBaselineCandidatePaired questions
Passed your review——0
Mean latency (ms)——0
Mean input + output tokens——0
Mean measured query cost (USD)——0

Human reviews are your judgments, not an automatic accuracy score. Each numeric row uses its own matched pairs. One-time indexing, storage, and platform fees are outside this per-query comparison.

Enter a question to enable export.