01 / Place the check after retrieval

Start with your existing retriever and a bounded set of candidate passages. A decision stage can inspect each question–passage pair before the answer model reads it. Preserve document IDs so accepted evidence can still be cited and rejected evidence can be inspected. This design adds a model call whose cost and delay must be measured.

TypeSafe · classifying RAG passages ↗

02 / Ask about relevance and support separately

A passage may discuss the right topic without supporting the requested answer. Define a relevance question and a support question. For multi-part requests, also check whether the retained set covers all required parts; one highly relevant passage does not prove completeness.

User question: Can I return an opened item after 20 days?
Passage A: Returns are accepted within 30 days.
Passage B: Opened items are excluded unless defective.

Checks:
- Is this passage relevant to the question?
- Does it contain a rule or exception needed for the answer?
- Does the retained evidence cover both timing and condition?

This is an authored example; no model probabilities are supplied.

03 / Measure the cost of missing evidence

Run the same questions with and without the decision stage. Keep the retriever, document snapshot and answer model fixed. Track evidence recall, citation support, final-answer quality, latency and cost. Inspect false rejections: deleting a short exception can make an answer cheaper and wrong.

Keep borderline passages or take a fallback retrieval path when the evidence is incomplete. Select thresholds on validation data and report performance on a separate test set.

04 / Keep each component's responsibility clear

An embedding model retrieves by representation similarity; a reranker orders candidates; a decision model can apply a question-specific criterion. These roles can overlap on a particular task, but a decision-model comparison should not silently treat their scores as the same quantity. Prefer the simpler pipeline if the extra stage does not improve the end-to-end result.