← All interview topics

FREE AI INTERVIEW Q&A / RAG & RETRIEVAL

RAG & retrieval interview questions

Ingestion, chunking, retrieval, grounding, and evaluation. Questions, answers, and explanations are presented in English, with Chinese source material translated and original question numbers preserved. Try each one before opening its matched answer and explanation.

100 free questions · English answers and source numbers · No sign-up

Showing 1–20 of 100 matching questions

Self-check progress: 0 of 100 reviewed · 0 marked “Got it”

QUESTION 01

Your RAG system retrieves the right chunk but the LLM still answers wrong. How do you debug this?

50 AI Engineer Interview Questions · Q1

Reveal source answer and explanation

Key concept: Can you isolate retrieval vs generation failures instead of guessing.

Reference answer: Separate retrieval from generation. First confirm the gold chunk is actually in the top-k passed to the model by logging the retrieved context. If it is there, the failure is generation: the model is ignoring context (lost-in-the-middle - put key chunks first or last), the prompt does not enforce grounding, or parametric knowledge is overriding it. Add "answer only from the context, else say not found," cut distractor chunks, and measure faithfulness with an NLI model or LLM-judge. If the chunk is NOT in top-k, it is retrieval: fix chunking, embeddings, or add a reranker.

QUESTION 02

Walk me through evaluating a RAG system - and why answer accuracy alone is not enough.

50 AI Engineer Interview Questions · Q2

Reveal source answer and explanation

Key concept: Do you decompose the pipeline and gate on evals, not vibes.

Reference answer: Evaluate the two stages separately. Retrieval: context recall (did we fetch the chunks needed) and context precision (how much noise). Generation: faithfulness (is the answer grounded, no hallucination) and answer relevance (does it address the query). A single end-to-end accuracy number hides which stage broke - you can get faithful answers on the wrong retrieved context. Build a golden set of 50-200 real queries with reference chunks and answers, run it in CI, and track all four metrics so a regression points at the culprit.

QUESTION 03

When would you NOT use RAG, and what would you use instead?

50 AI Engineer Interview Questions · Q3

Reveal source answer and explanation

Key concept: Do you treat RAG as one option with costs, not a default.

Reference answer: RAG adds latency (retrieval + rerank + a bigger prompt) and a retrieval failure mode. Skip it when knowledge is small and static (put it in the system prompt or fine-tune), when the task is reasoning or transformation rather than fact lookup, or when you need sub-100ms latency and cannot afford the hop. Alternatives: long-context stuffing for a bounded corpus, fine-tuning for behaviour and format, or a plain API/SQL call when the "knowledge" is structured data best fetched by a query, not similarity search.

QUESTION 04

How do you keep a RAG index fresh when source documents change constantly?

50 AI Engineer Interview Questions · Q8

Reveal source answer and explanation

Key concept: Do you think about the data pipeline and staleness, not a one-time build.

Reference answer: Treat the index as a derived store with an invalidation pipeline. Track document versions or hashes; on change, re-chunk and re-embed only the affected docs (upsert by stable chunk IDs) instead of rebuilding everything. For deletes, tombstone immediately so removed content cannot be retrieved. Monitor ingestion lag so you know how stale retrieval can be, and store an as-of timestamp so answers can cite recency. Full re-index only when the embedding model or chunking strategy changes - the one case you cannot update incrementally.

QUESTION 05

Cross-encoder reranking is accurate but slow. How do you use it without blowing the latency budget?

50 AI Engineer Interview Questions · Q10

Reveal source answer and explanation

Key concept: Do you balance retrieval quality against a real latency budget.

Reference answer: Reranking costs one forward pass per candidate, so never rerank the corpus - retrieve a broad top-50/100 cheaply, then rerank to a top-3/5. Control cost with a smaller or distilled reranker, batching, and caching scores for repeat query-doc pairs. If p95 is still tight, make reranking conditional (skip when top vector scores are already well separated) or use a late-interaction model (ColBERT) that is cheaper than full cross-encoding. Measure the recall gain against the 50-200ms it adds; if it is not moving answer quality, cut it.

QUESTION 06

How should the retrieval efficiency of a vector database be evaluated?

AI Agent Development: 158 Interview Questions · 6.1.1. · p. 39

Reveal source answer and explanation

Key concept: Retrieval performance evaluation metrics, such as query speed, throughput, and latency.

Explanation: By setting up standard test scenarios, collect the query time and response speed of different databases under the same hardware and data scale, and compare their efficiency in index construction and querying. It is necessary to consider the database's index structure (such as Annoy, FAISS, HNSW, etc.) as well as hardware resources (such as GPU, SSD). Be careful to ensure that the test environment is consistent to avoid performance bias. The impact of data dimensionality on efficiency and performance under concurrent queries should also be considered. A common mistake is focusing only on single-query time and ignoring overall throughput or index maintenance efficiency.

Reference answer: Evaluating the retrieval efficiency of a vector database should combine query speed, throughput, and latency metrics, use unified test data and hardware environments for comparison, and also consider the impact of index structure and hardware configuration on performance.

QUESTION 07

What factors should be considered when choosing a vector database?

AI Agent Development: 158 Interview Questions · 6.1.4. · p. 41

Reveal source answer and explanation

Key concept: Evaluation criteria include performance, scalability, feature support, and ecosystem.

Explanation: During evaluation, consider the database's retrieval speed and accuracy to meet business needs; the data scale and growth rate, which determine whether horizontal scaling and distributed storage are supported; the index algorithms and supported similarity measurement methods, and whether they meet specific application scenarios (such as text, images, etc.); and whether online updates and deletions are supported and their impact on performance. In addition, depending on the integration environment, consider the database's API friendliness, compatibility, and maintenance difficulty. Community activity and ecosystem support should also be observed to reduce later maintenance costs. A common mistake is underestimating the complexity of future expansion and maintenance while focusing only on current performance metrics.

Reference answer: When choosing a vector database, factors such as performance, scalability, index and query support, data type compatibility, maintenance difficulty, and ecosystem support should be comprehensively considered to ensure that it meets the long-term development needs of the business.

QUESTION 08

How do embedding models improve RAG retrieval efficiency?

AI Agent Development: 158 Interview Questions · 6.2.1. · p. 41

Reveal source answer and explanation

Key concept: Understand the role of embedding models in information retrieval and their difference from traditional keyword matching.

Explanation: Considering that traditional retrieval struggles with semantic matching, embedding models convert text into vectors and use similarity in vector space for retrieval, thereby achieving semantic-level matching. It is necessary to consider the model's vector representation capability and retrieval speed, and choose suitable embedding algorithms (such as BERT, SBERT) and index structures (such as FAISS) to improve efficiency. Also pay attention to the balance between vector dimensions and storage space to avoid the computational burden caused by excessively high dimensions, while focusing on the match between the model's semantic representation and actual retrieval needs.

Reference answer: Embedding models convert text into high-dimensional vectors to enable similarity computation in semantic space, speeding up retrieval and improving matching accuracy. This mechanism plays a key role in context recall within RAG systems, helping retrieve relevant and rich textual information and reducing the false detection rate.

QUESTION 09

How can the accuracy of embedding retrieval be improved?

AI Agent Development: 158 Interview Questions · 6.2.3. · p. 42

Reveal source answer and explanation

Key concept: The methods and strategies for improving the effectiveness of vector similarity retrieval.

Explanation: This can be approached from two aspects: data preparation and model training. First, use high-quality, rich, and diverse training data for fine-tuning to enhance the model's semantic understanding capability. Second, apply fine-tuning techniques targeted at specific tasks so that the resulting vectors better meet business needs. In addition, optimize the index structure (such as using efficient indexing algorithms like IVF or HNSW) to improve retrieval precision and speed. In post-processing, adjust the similarity threshold to exclude low-relevance results, and optimize results by combining rule-based filtering measures. Also pay attention to the consistency of embedding vector scale to avoid numerical bias affecting similarity computation.

Reference answer: By fine-tuning the model, optimizing the index structure, and adjusting the similarity threshold, the accuracy of Embed-ding retrieval can be effectively improved, ensuring the relevance and precision of retrieval results.

QUESTION 10

How can the diversity and relevance of retrieval recall be balanced?

AI Agent Development: 158 Interview Questions · 6.3.1. · p. 43

Reveal source answer and explanation

Key concept: The trade-off between diversity and relevance in retrieval strategies, and optimizing recall effectiveness.

Explanation: Analyze that in RAG, the diversity of recall results can avoid bias from a single information source but may sacrifice some relevance; emphasizing relevance may limit diversity, leading to information limitations. Strategies need to be adjusted according to task objectives, such as introducing diversity re-ranking, multimodal retrieval, or achieving dynamic balance by adjusting retrieval weights. At the same time, consider user needs to ensure that neither relevance nor diversity is overemphasized, avoiding impact on the final context matching effect. Considering means such as parameter tuning in retrieval algorithms and the fusion of different recall strategies is also an important method for ensuring balance.

Reference answer: By combining diversity re-ranking techniques, such as favoring diversity or introducing diversity metrics, and adjusting the relevance and diversity trade-off parameters in the retrieval model, ensure that retrieval results both cover diverse information and maintain a certain level of relevance, thereby improving the overall performance of the RAG system.

QUESTION 11

How can retrieval strategies be optimized in multimodal retrieval?

AI Agent Development: 158 Interview Questions · 6.3.3. · p. 44

Reveal source answer and explanation

Key concept: The design and optimization of retrieval strategies for multimodal information fusion.

Explanation: Multimodal retrieval involves multiple data types such as text, images, and videos, and optimization strategies need to consider the characteristics and contribution proportions of different modalities. Fusion models, such as a multimodal embedding space, can be used to map information from different modalities into a unified representation, improving retrieval efficiency and accuracy. In retrieval strategies, pay attention to adjusting the weights of information from different modalities, as different modalities may vary in importance across scenarios. Also consider the complementarity and redundancy between modalities, and design reasonable fusion mechanisms (such as weighted fusion, attention mechanisms, etc.) to improve the overall diversity and relevance of retrieval. At the same time, also consider the noise and inconsistency of multimodal data, and adopt robust filtering and optimization algorithms.

Reference answer: By establishing a multimodal embedding space, adopting a fusion mechanism that dynamically adjusts the weights of different modalities, and introducing attention mechanisms, optimize the integration of multimodal information and retrieval strategies to improve the accuracy and robustness of multimodal retrieval.

QUESTION 12

How can an effective re-ranking algorithm be designed?

AI Agent Development: 158 Interview Questions · 6.4.1. · p. 44

Reveal source answer and explanation

Key concept: Understand the design principles and optimization strategies of re-ranking algorithms, especially the re-ranking process for retrieval rankings in RAG.

Explanation: Analyze the algorithm objectives to ensure that re-ranking can enhance relevance or diversity. Consider feature selection, model selection (such as ranking models or deep models), feature engineering, and model training methods (such as pointwise learning, structured learning). Efficiency issues also need to be considered to avoid slowing down the overall response speed. During the design process, pay attention to data bias and overfitting risks, and perform cross-validation and parameter tuning when necessary. Possible pitfalls include poor real-time performance caused by excessive model complexity and bias caused by over-reliance on a single feature.

Reference answer: Designing an effective re-ranking algorithm should combine feature engineering, model selection, and efficiency optimization to ensure that ranking results improve relevance and diversity while also having good real-time performance. Common methods include learning-to-rank (such as RankNet, LambdaRank) or strongly adjusted model tuning, fusing multi-source information, and enhancing robustness.

QUESTION 13

How can the effectiveness of a re-ranking model be evaluated?

AI Agent Development: 158 Interview Questions · 6.4.2. · p. 45

Reveal source answer and explanation

Key concept: Master evaluation metrics and experimental design, and understand multidimensional methods for measuring model performance.

Explanation: Consider commonly used ranking effectiveness metrics, such as NDCG, MRR, MAP, etc., and choose evaluation metrics suitable for the application scenario. When designing offline evaluation, a test set with real annotations is needed to ensure the representativeness of the data. A/B testing or online experiments can be conducted to verify the model's improvement in user experience. During the evaluation process, pay attention to the sensitivity and bias of the metrics to ensure their stability over time. Attribution analysis and error analysis are also part of optimization, helping identify the model's weaknesses and potential directions for improvement.

Reference answer: The effectiveness of a re-ranking model is mainly evaluated through metrics such as NDCG, MRR, and MAP, combined with offline testing and online A/B testing, to ensure that while improving ranking quality, the model also guarantees user experience and system stability.

QUESTION 14

What retrieval methods are used in RAG? What are their advantages and disadvantages?

RAG (Retrieval-Augmented Generation): 115 Interview Questions · 1.1.3. · p. 6

Reveal source answer and explanation

Key concept: Master the types of different retrieval techniques and their respective applicable scenarios.

Explanation: Analyze commonly used retrieval methods, including traditional keyword-based retrieval (such as BM25) and vector-based dense retrieval (such as DPR and FAISS). Keyword retrieval is simple and fast, but its semantic understanding is not strong and it is heavily affected by query terms. Vector retrieval can capture semantic relationships and support more flexible matching, but its computational cost is higher. When considering implementation complexity, scalability, and scenario requirements, it is necessary to weigh and choose an appropriate retrieval strategy, recognizing that advantages and disadvantages differ across scenarios and that it is easy to fall into the problem of focusing only on algorithms while ignoring practical application.

Reference answer: The main retrieval methods include traditional keyword-based retrieval (such as BM25) and dense vector retrieval (such as DPR and FAISS). The advantage of keyword retrieval is that it is simple to implement and fast, but its semantic understanding is limited; vector retrieval can better capture semantic relationships and is suitable for complex semantic matching, but it has high computational resource requirements. In practical applications, the choice must be made according to the scenario, and the two can be combined to achieve optimized results.

QUESTION 15

How does a RAG model achieve dynamic knowledge updates?

RAG (Retrieval-Augmented Generation): 115 Interview Questions · 1.1.4. · p. 7

Reveal source answer and explanation

Key concept: Understand the dynamic maintenance of the knowledge base in a RAG system and the model's adaptability.

Explanation: Consider that the knowledge base of RAG (such as the retrieval index) is one of the core parts of the model. Achieving dynamic updates usually involves continuously adding new information to the index or knowledge base to ensure that the retriever can obtain the latest content. Technically, incremental index updates or real-time index refresh can be used, and it is also necessary to ensure synchronization between the retrieval and generation parts. Attention should also be paid to how to reduce retrieval latency and avoid affecting the overall response speed. An easy mistake is to ignore the synchronization mechanism of the index and the impact of updates on system performance.

Reference answer: By incrementally building indexes or using dynamic index management strategies, new knowledge is incorporated into the index to ensure the retrieval results are up to date. In system design, periodic refresh or real-time update mechanisms can be set up so that the knowledge base stays in sync with actual information. This approach supports RAG in continuously providing accurate answers in a constantly changing knowledge environment, while it is also necessary to ensure that the update process does not significantly affect retrieval and generation efficiency.

QUESTION 16

What is the core difference between RAG models and traditional generative models?

RAG (Retrieval-Augmented Generation): 115 Interview Questions · 1.2.1. · p. 7

Reveal source answer and explanation

Key concept: Understand the difference between RAG and traditional generative models in information retrieval and generation methods.

Explanation: Analyze that traditional generative models rely on end-to-end training and generate output directly from input, while RAG combines retrieval and generation, integrating the retrieved context into the generation process. This involves understanding RAG's retrieval-generation architecture and identifying its advantage of introducing an external knowledge base. When thinking about this, note that RAG should not be equated with a pure retrieval or generation model, but is an architecture combining both; an easy mistake is to confuse the two modes of operation.

Reference answer: RAG models enhance the factuality and controllability of content by retrieving relevant information before generating content; whereas traditional generative models rely entirely on training data for end-to-end learning and lack the dynamic introduction of external knowledge.

QUESTION 17

How does a RAG model combine the two parts of retrieval and generation?

RAG (Retrieval-Augmented Generation): 115 Interview Questions · 1.2.4. · p. 8

Reveal source answer and explanation

Key concept: Understand the overall workflow and structural design of the RAG model.

Explanation: Analyze the RAG process: first, use the retrieval module to find relevant context from the knowledge base, then pass the retrieval results as part of the input to the generative model for content generation. Consider the coupling method of the two, including the encoding of retrieval results, the fusion strategy, and how to ensure that the generated content fully utilizes the retrieved information and avoids information omission or bias. Pay attention to the interface design details between retrieval and generation to avoid information silos or leakage.

Reference answer: A RAG model first uses a retrieval module to find relevant information, then inputs the retrieved content as context into the generative model to jointly generate an answer, achieving dynamic fusion of information.

QUESTION 18

Why is normalization performed in vector retrieval?

RAG (Retrieval-Augmented Generation): 115 Interview Questions · 2.1.1. · p. 9

Reveal source answer and explanation

Key concept: Understand the role and necessity of vector normalization, especially as the basis of cosine similarity.

Explanation: In vector retrieval, normalization can make vectors of different scales comparable and avoid the impact of length differences. Especially when using cosine similarity, the direction of the vector is more important than its length, and normalization ensures that similarity depends only on angle rather than magnitude. Conversely, if not normalized, long vectors may affect distance calculation and cause similarity bias. Considering different retrieval scenarios, whether normalization is needed should also be combined with the similarity metric used.

Reference answer: Normalization is performed to ensure that vectors rely only on direction to calculate similarity. Especially when using cosine similarity, after standardization the vector's magnitude is 1, which helps improve retrieval accuracy and stability. This is also a commonly used preprocessing step in vector retrieval.

QUESTION 19

What are common methods for vector quantization?

RAG (Retrieval-Augmented Generation): 115 Interview Questions · 2.1.5. · p. 11

Reveal source answer and explanation

Key concept: Master vector quantization methods and their applicable scenarios.

Explanation: Vector quantization reduces storage and computational burden by mapping vectors in continuous space to representatives in a finite set. Common methods include K-means clustering (creating multiple cluster centroids as representatives), Product Quantization (PQ, splitting vectors into subvectors and quantizing them separately), Optimized Vector Quantization, etc. Each method has its own advantages and limitations in compression rate, computational complexity, and retrieval accuracy. In design, the requirements of the retrieval task, hardware limitations, and maintenance costs should be considered. In addition, information loss caused by quantization should be avoided from affecting retrieval accuracy.

Reference answer: Commonly used vector quantization methods include K-means, Product Quantization, etc. They reduce storage space through compression while maintaining relatively good retrieval performance, and are widely used in large-scale vector retrieval tasks.

QUESTION 20

How should a suitable retriever be selected to adapt to different tasks?

RAG (Retrieval-Augmented Generation): 115 Interview Questions · 2.2.1. · p. 12

Reveal source answer and explanation

Key concept: Principles for selecting retriever types and applicable scenarios.

Explanation: This examines whether the candidate understands the characteristics and applicable scenarios of different retrievers (such as inverted index-based, vector retrieval, and hybrid retrieval). The characteristics of task requirements should be analyzed, such as keyword matching versus semantic understanding, real-time requirements, data scale, etc. Evaluate whether the candidate can reasonably match retrieval technologies by combining task complexity and system performance. Note that it is necessary to avoid seeing only the technology while ignoring limitations in the actual application environment, such as storage cost and computational resources. Considering easily confused points, the candidate should be clear about the trade-offs between accuracy and efficiency of different retrievers.

Reference answer: Selecting a suitable retriever should be based on task requirements. For keyword matching tasks, inverted indexes are highly efficient; for semantic search needs, vector retrieval is better; hybrid methods can combine the advantages of both. In addition, considering data scale, update frequency, and real-time requirements, the retrieval strategy should be reasonably adjusted to ensure a balance between effectiveness and performance.

Your self-check is kept only on this page and resets when you leave. It is not an automated score or hiring prediction.

Explore the underlying concepts

The study-bank title and original question number appear on every question. The links below are additional technical reading. Source answers are study references; check version-specific claims against current documentation.

LangChain: Retrieval ↗Elasticsearch: Reciprocal rank fusion ↗