STEP 01
Check whether useful text exists
A PDF can display perfectly while providing little usable text to an index. Your first checkpoint is the extracted content, not the final chatbot answer.
- Compare extracted text against a representative sample of original pages, including a dense page and a table.
- If extraction is empty or scrambled, run a suitable OCR/document parser and inspect its output before indexing.
- Check decimals, negative signs, units, column headers, and reading order. Do not rely on a visual preview alone.
STEP 02
Preserve what an answer needs
Write down the fields needed for the questions you expect. The goal is readable evidence with traceable provenance.
- Keep document titles, section labels, and page references where available.
- Repeat table headers alongside row values in your prepared text if the extraction loses those relationships.
- Do not enable cleanup that removes URLs or email addresses when those values are themselves answers users need.
STEP 03
Verify extraction, then retrieval
Choose a distinctive sentence and a number from the original. Both must survive the preparation pipeline.
- Find them in the imported content and inspect their chunks after indexing.
- Run an exact question and a paraphrase through search alone, then inspect a complete app run.
- If the underlying text is still wrong, return to extraction. Changing retrieval weights cannot reconstruct absent source facts.
Give your changes a fair test.
Open comparison worksheet Use the same questions, inspect the evidence, and record what changed.