KNOWLEDGE LAB / PIPELINE REFERENCE

Explore the pipeline, stage by stage.

Compare preparation and retrieval options, read their practical lessons, and study the settings. These choices create a reference checklist.

Start with an application scenario →

01 / MAP YOUR PIPELINE

Different sources. One learning path.

Explore a scenario, then build a route around your own documents.

My first chatbot

One file, a few known answers. Read the lesson, or use its recommended settings as a starting point.

Read this lesson

Your pipeline

Based on My first chatbot · Recommended settings. Changes here leave the scenario above unchanged.

01Data sources

Select all that apply. Keep at least one source.

02Extraction for each source

Choose every kind of content you need to read. Each selected method has its own lesson; apply it only to the matching documents.

Local files

Select one or more methods for this source.

After extraction, these sources feed one knowledge base. Its chunk structure, index, and retrieval strategy are shared.

03Processing

How should it be split? Choose one.

04Storage & index

How will it be indexed? Choose one.

05Retrieval

How will you find it? Choose one.

Your selections are saved in this page’s URL. Choosing a scenario only previews its lesson.

Your sources → shared knowledge base
  • Local filesReadable text

General chunks Vector + text Hybrid

02 / LEARN & BUILD

Work through every source. Then bring them together.

Your plan has 5 steps. Prepare each source and its extraction methods, then configure the shared knowledge base.

STEP 1 OF 5 / DATA SOURCE

Local files

Start with a small collection whose answers you can verify yourself.

  1. Choose one current file and give it a recognizable name.
  2. Write one question and mark the passage that answers it.
  3. Follow the first-answer lesson; parse the file and record its stable source ID in your ingestion job.
START WITH

harbor-studio-v1.txt

LOOK FOR

Source record: file name + version + original text

Illustrative example · Fictional studio facts · Not a live pipeline run
Before moving on

You know which file, version, and passage should answer your first question.

Save this route and its checklist

Includes every selected source, its extraction steps, compatibility notes, and shared tuning guidance. Your page URL also preserves the selected options.

03 / TUNE YOUR CHOSEN ROUTE

One change. A clear comparison.

Build and check your baseline first. These controls follow the choices above.

01General chunks

Keep the right context

Delimiter, length & overlap

Preview the longest answer. Adjust a boundary if it separates a fact from its exception.

How to check the change

The complete supporting fact survives in the returned context.

Build a tiny RAG with one file
02Vector + text

Check the index

Embedding model & source metadata

Record the embedding model and document version; check indexing errors before tuning search.

How to check the change

Known evidence is indexed and traceable. Record actual usage rather than assuming a cost saving.

Find a fact your RAG keeps missing
03Hybrid

Find enough evidence

Weights, Top K & threshold

Try the Top K lesson first. Keep weights and threshold fixed while comparing the returned passages.

How to check the change

Inspect actual passages. Save chosen settings in your search service and answer builder.

Change just one setting: Top K
04Filters · ranking · citations

Refine and verify

Source scope and evidence quality

Use metadata to keep the intended version in scope. If ranking is the problem, compare reranking separately and measure its added cost and latency.

How to check the change

Check the cited passage, a covered question, and an uncovered question in a fresh conversation.

Show the source behind an answer

Changing chunk structure requires reindexing. Compare a different structure in a separate test index. See Technical reference ↗.

04 / VERIFY THE WHOLE PATH

Does the answer hold up?

Connect the knowledge base to your app. For each source, test a known fact and its citation. Then ask a question outside the collection.

After retrieval: check the evidence.

Study how decision models can check relevance and coverage, including the evidence a filter removes.

Read the evidence-checking guide →