03 / VERIFY
Evaluate your RAG changes.
Use a simple comparison worksheet to capture the evidence. Run the same questions against two configurations in your own workspace.
Run tests in your own RAG app, then enter the results here. Nothing in this worksheet is uploaded or automatically scored. Download your CSV before leaving; page entries are not saved. Read the testing guide first →
1. Record what changed
Keep documents, models, and prompts fixed unless they are the change you are testing.
Use the same currency and billing scope for both runs. Currency selection labels your entries; it does not convert values. Leave unknown measurements blank.
2. Compare the same questions
Include answerable and unanswerable questions. Reserve some for validation.
3. Review paired results
Only questions with a name and measurements on both sides count.
| Measure | Baseline | Candidate | Paired questions |
|---|---|---|---|
| Passed your review | — | — | 0 |
| Mean latency (ms) | — | — | 0 |
| Mean input + output tokens | — | — | 0 |
| Mean measured query cost (USD) | — | — | 0 |
Human reviews are your judgments, not an automatic accuracy score. Each numeric row uses its own matched pairs. One-time indexing, storage, and platform fees are outside this per-query comparison.
Enter a question to enable export.