STEP 01
Give each model a specific job
The local starter uses keyword overlap and needs no embedding model. In a semantic route, an embedding model represents documents and queries as vectors; the answer model writes a response from retrieved text. An optional reranker changes candidate order. These are separate decisions.
- First read the starter's printer passage. If the fact is missing from the source or extraction, choosing a larger answer model will not repair that input.
- Compare the literal word ‘printer’ with ‘Where can I print something?’ to design an embedding evaluation. Include your actual document languages and terminology.
- Give two answer-model candidates the same passages and question. Check the second-floor location, blue cabinet, and supporting citation before comparing measured usage.
Your notes distinguish finding evidence, ranking it, and writing a supported answer.
STEP 02
Record an embedding configuration you can reproduce
Choose a candidate using language coverage, input limits, access requirements, and tests on your questions. Record the exact model ID, any document/query encoding instructions, and output dimensions. A matching dimension alone does not make two embedding models compatible.
- Use the selected model's documented document and query encoding modes. Keep both sides in its compatible embedding space.
- Match the database vector field to the actual output dimension and choose a supported distance metric using the embedding and database documentation.
- When changing the model, dimensions, or preprocessing, build a separate test index and re-embed the documents. Keep the earlier baseline available for comparison.
Embedding model + revision: not selected
Document/query encoding modes: not recorded
Output dimensions: not measured
Document languages: English + your actual languages
Answer model + revision: not selected
Prompt revision: baseline-1
Evidence: printer, booking, and uncovered parking questionsIf it does not work
If a vector is rejected, check its length against the schema. If it is accepted but retrieval becomes poor after a model change, check model identity and preprocessing too.
STEP 03
Set the answer boundary before tuning generation
Define the source contract first: answer only from supplied passages, cite returned IDs, and identify missing evidence. Temperature is not a factuality switch, and available generation controls vary by model.
- Choose a response format: concise answer, source IDs, and an explicit insufficient-evidence outcome. Validate that cited IDs were retrieved and that their text supports the claim.
- Measure the actual evidence tokens. Set an output allowance for the required answer format and leave room for instructions and the question within the model's context limit.
- Keep supported generation settings fixed during the initial comparison. If the chosen model exposes temperature, change it in a separate experiment; do not assume every model accepts it.
- Run the printer question and the parking question together. A fluent invented parking policy fails even when the printer answer succeeds.
A written model choice backed by paired examples; unknown latency, cost, and quality remain unmeasured.
Use the same questions, inspect the evidence, and record what changed.