All knowledge base guides

MODELS & SETTINGS / FIELD GUIDE

Choose embedding and answer models separately

Use the printer and parking exercises to separate retrieval failures from answer-model failures, then record a reproducible model choice.

9 min read · SaveMyToken editorial · Documentation reviewed 2026-10-05 · Build in your own stack

STEP 01

Give each model a specific job

The local starter uses keyword overlap and needs no embedding model. In a semantic route, an embedding model represents documents and queries as vectors; the answer model writes a response from retrieved text. An optional reranker changes candidate order. These are separate decisions.

  1. First read the starter's printer passage. If the fact is missing from the source or extraction, choosing a larger answer model will not repair that input.
  2. Compare the literal word ‘printer’ with ‘Where can I print something?’ to design an embedding evaluation. Include your actual document languages and terminology.
  3. Give two answer-model candidates the same passages and question. Check the second-floor location, blue cabinet, and supporting citation before comparing measured usage.
You should see

Your notes distinguish finding evidence, ranking it, and writing a supported answer.

STEP 02

Record an embedding configuration you can reproduce

Choose a candidate using language coverage, input limits, access requirements, and tests on your questions. Record the exact model ID, any document/query encoding instructions, and output dimensions. A matching dimension alone does not make two embedding models compatible.

  1. Use the selected model's documented document and query encoding modes. Keep both sides in its compatible embedding space.
  2. Match the database vector field to the actual output dimension and choose a supported distance metric using the embedding and database documentation.
  3. When changing the model, dimensions, or preprocessing, build a separate test index and re-embed the documents. Keep the earlier baseline available for comparison.
Model selection notebook — fill in after checking your provider
Embedding model + revision: not selected
Document/query encoding modes: not recorded
Output dimensions: not measured
Document languages: English + your actual languages
Answer model + revision: not selected
Prompt revision: baseline-1
Evidence: printer, booking, and uncovered parking questions

If it does not work

If a vector is rejected, check its length against the schema. If it is accepted but retrieval becomes poor after a model change, check model identity and preprocessing too.

STEP 03

Set the answer boundary before tuning generation

Define the source contract first: answer only from supplied passages, cite returned IDs, and identify missing evidence. Temperature is not a factuality switch, and available generation controls vary by model.

  1. Choose a response format: concise answer, source IDs, and an explicit insufficient-evidence outcome. Validate that cited IDs were retrieved and that their text supports the claim.
  2. Measure the actual evidence tokens. Set an output allowance for the required answer format and leave room for instructions and the question within the model's context limit.
  3. Keep supported generation settings fixed during the initial comparison. If the chosen model exposes temperature, change it in a separate experiment; do not assume every model accepts it.
  4. Run the printer question and the parking question together. A fluent invented parking policy fails even when the printer answer succeeds.
You should see

A written model choice backed by paired examples; unknown latency, cost, and quality remain unmeasured.

Give your changes a fair test.

Use the same questions, inspect the evidence, and record what changed.

Open comparison worksheet