Prepare for AI engineer and Forward Deployed Engineer (FDE) interviews with free topic questions, matched source answers, foundation practice, and customer scenarios. All questions and answers are in English, with original source numbers preserved.
A policy assistant cites the correct handbook but tells a customer that opened products can be returned. The handbook contains an exception excluding them. How would you isolate the failure?
Reveal an answer outline
Clarify first
Was the exception in the exact context sent to the model, or only somewhere in the source document?
Inspect the actual evidence path. Reproduce the case with a redacted trace. Follow the exception through extraction, chunking, candidate retrieval, reranking, and final context assembly. A correct document ID does not establish that the decisive sentence reached generation.
Test the failing stage. If the exception is absent, inspect chunk boundaries and candidate coverage. If present, test whether the answer preserves both the rule and exception with reduced distracting context. Change one factor at a time and retain the original failure as a regression case.
Judge claims, not citation presence. Check whether each material claim follows from the accessible passage. Include negative examples with valid citations but false statements. Review abstention when the policy cannot support an answer.
The tradeoff: More retrieved text can recover an exception but can also add conflicting policy versions. Measure evidence completeness and answer correctness together.
A maintenance handbook mixes paragraphs, equipment tables, and footnotes. Fixed-size chunks separate a replacement interval from its equipment heading. How would you choose and test a better chunking strategy?
Reveal an answer outline
Clarify first
Which queries require multiple cells, a heading, or a footnote to be interpreted correctly?
Define an evidence unit. Identify complete units for the target queries: a procedure with prerequisites, or a table row with headers and applicable footnotes. Preserve source locations and document versions so a reviewer can inspect the original context.
Compare concrete candidates. Try structure-aware boundaries and bounded parent-context expansion against the current splitter. Avoid treating one token length or overlap percentage as universally correct. Keep extraction, embeddings, and evaluation queries fixed during the first comparison.
Evaluate downstream use. Label whether each retrieved result carries enough information to answer. Then inspect generated answers, context size, and latency. Include long tables and exception-heavy procedures in a held-out set before selecting the new default.
The tradeoff: Larger evidence units retain meaning but may bring irrelevant rows. Smaller searchable units plus controlled surrounding context are another option to test.
A repair assistant handles both “E-417 on ZX-8” and “the pump rattles after startup.” Vector-only retrieval handles some paraphrases but misses exact code matches. How would you improve and evaluate retrieval?
Reveal an answer outline
Clarify first
Do extraction and indexing preserve punctuation, model numbers, and code-to-product relationships?
Build query slices. Create labeled examples for exact identifiers, paraphrases, and mixed queries. Inspect failures caused by broken text extraction or normalization before adding another retrieval layer.
Combine eligible candidates. Compare lexical and dense retrieval, then combine their ranked lists with an explicit method such as reciprocal rank fusion. Apply the same access scope to both paths. Deduplicate by stable passage IDs before assembling evidence.
Test optional reranking. Evaluate a reranker over a bounded candidate set only if the combined ranking leaves useful evidence buried. Track per-slice coverage, downstream answers, and latency; keep a simple baseline to show which stage earns its cost.
The tradeoff: Hybrid retrieval introduces another index and tuning surface. It is useful only if the measured coverage gains justify its operating and query costs.
HR replaces a leave policy and withdraws the old document. The assistant still returns the old allowance, sometimes with a working citation. Design an update and deletion process that you can verify.
Reveal an answer outline
Clarify first
Is the stale result coming from the source, the index, a response cache, or conversation history?
Trace identity and versions. Give documents and chunks stable identities, track revisions, and record effective dates separately from ingestion times. Inspect the answer's source version and all caches involved in producing it.
Make withdrawal observable. Immediately exclude withdrawn records from eligible retrieval. Propagate deletion or replacement to indexes and dependent caches, and reconcile failed updates. A user continuing an old conversation should be rechecked against the current authorized source state.
Verify with a known question. Run a regression question that distinguishes the two policies. Check retrieved IDs, citations, cache behavior, and the final answer. Monitor update lag and expose uncertainty if a required source is temporarily unavailable.
The tradeoff: Immediate exclusion can reduce answer coverage while replacement indexing completes. Returning an explicit gap is preferable to silently treating withdrawn policy as current.