All knowledge base guides

DOCUMENT PREPARATION / FIELD GUIDE

Fix scanned PDFs before tuning RAG

Verify OCR, reading order, and table structure before spending time on retrieval settings.

5 min read · SaveMyToken editorial · Documentation reviewed 2026-10-05 · Build in your own stack

STEP 01

Check whether useful text exists

A PDF can display perfectly while providing little usable text to an index. Your first checkpoint is the extracted content, not the final chatbot answer.

  1. Compare extracted text against a representative sample of original pages, including a dense page and a table.
  2. If extraction is empty or scrambled, run a suitable OCR/document parser and inspect its output before indexing.
  3. Check decimals, negative signs, units, column headers, and reading order. Do not rely on a visual preview alone.

STEP 02

Preserve what an answer needs

Write down the fields needed for the questions you expect. The goal is readable evidence with traceable provenance.

  1. Keep document titles, section labels, and page references where available.
  2. Repeat table headers alongside row values in your prepared text if the extraction loses those relationships.
  3. Do not enable cleanup that removes URLs or email addresses when those values are themselves answers users need.

STEP 03

Verify extraction, then retrieval

Choose a distinctive sentence and a number from the original. Both must survive the preparation pipeline.

  1. Find them in the imported content and inspect their chunks after indexing.
  2. Run an exact question and a paraphrase through search alone, then inspect a complete app run.
  3. If the underlying text is still wrong, return to extraction. Changing retrieval weights cannot reconstruct absent source facts.
Give your changes a fair test.

Use the same questions, inspect the evidence, and record what changed.

Open comparison worksheet