All knowledge base guides

STORAGE & INDEXING / FIELD GUIDE

Choose a RAG database and understand its settings

Compare a local baseline, PostgreSQL with pgvector, and a dedicated vector service using source records, version updates, and access-filter exercises.

10 min read · SaveMyToken editorial · Documentation reviewed 2026-10-05 · Build in your own stack

STEP 01

Choose from your constraints, not a popularity list

For the one-file lesson, the downloadable keyword retriever is enough to inspect evidence. When evaluating a persistent vector store, write down your existing stack, document volume, update rate, access rules, and who will operate it. The following are options to investigate, not a benchmark ranking.

  1. Already using PostgreSQL? Evaluate pgvector alongside your existing records and SQL workflow. Confirm the extension is available in your environment and record PostgreSQL and pgvector versions.
  2. Need a separate vector service? Evaluate Qdrant's collection and payload model against the same ingestion, filtering, and update exercises. Record server and client versions.
  3. Prefer a managed service? Check data location, backup/restore, access controls, supported search modes, current limits, and current pricing in the chosen provider's documentation. Managed hosting is an operational choice, not proof of better retrieval.
You should see

A short selection note explaining your constraints and the tests a candidate must pass.

STEP 02

Map a passage to a stored record

Keep readable text and provenance alongside vectors. In pgvector, a vector column can declare its dimensions and an index's operator class corresponds to its distance metric. A Qdrant vector configuration declares size and distance; named vectors can have separate configurations. Copy settings only after checking the chosen embedding model.

  1. Record the embedding model, actual output dimensions, and supported metric together. Do not paste a dimension from an unrelated tutorial.
  2. Keep document_id, version, passage_id, title, source location, and the text needed to inspect a match. Add the fields used by your application's access and version filters.
  3. Start with correctness checks before tuning approximate-index options. In pgvector, compare an approximate index against exact search; pay attention to how filters affect returned results.
Illustrative passage record — field names are application choices
document_id: harbor-studio
version: v1
passage_id: harbor-studio-v1.txt#3
text: [copy the complete booking paragraph]
source: harbor-studio-v1.txt
access_scope: studio-members
embedding_model: [exact selected model ID]
embedding_dimensions: [actual vector length]

If it does not work

Metadata does not enforce permissions by itself. Your retrieval path must apply the authenticated user's allowed scope before passages reach the model.

STEP 03

Use the existing version-change exercise as an acceptance test

The fictional Harbor Studio FAQ changes its booking limit from 60 minutes in v1 to 90 minutes in v2. Use this existing exercise to examine storage behavior instead of declaring success after creating a collection.

  1. Record which version is current and how replaced passages are removed or excluded from ordinary queries. Keep historical versions separately if your use case needs them.
  2. After the update, ask the booking question. Inspect both the returned text and version; changing a title without replacing old passages is insufficient.
  3. Design a second test using two access scopes. Verify an unauthorized scope receives no protected passage, including when an otherwise matching query is used.
  4. Record query latency, index size, and update behavior only after measuring them in your own environment.
You should see

You can explain where the current answer is stored, how it is replaced, and how access is restricted.

Give your changes a fair test.

Use the same questions, inspect the evidence, and record what changed.

Open comparison worksheet