Adding retrieved documents to a language model raises several engineering questions at once. The system must find useful evidence within a budget, handle hostile or biased material, respond to changing user needs and preserve evidence across reasoning steps. A new survey organizes retrieval-augmented generation around those four concerns.
Mapping the RAG Landscape proposes axes for efficiency, defense, interactivity and complex reasoning. It is a literature synthesis and taxonomy, rather than a new model with a measured accuracy improvement. Its value is in making design trade-offs easier to compare across otherwise different RAG methods.
Four questions for system builders
The efficiency axis examines how retrieval and compression can maintain factual accuracy under inference-time constraints. Returning more context can increase cost without solving the problem of finding the right evidence. The survey treats retrieval scoring and context construction as decisions with an accuracy-cost trade-off.
Defense covers vulnerabilities introduced by the external corpus itself, including poisoning, demographic bias amplification and privacy leakage. Grounding an answer in retrieved text does not make that text trustworthy. A system's choice of sources and filtering has to accompany its relevance ranking.
The interaction axis asks how retrieval can account for user intent, history and preferences while retaining factual precision. The reasoning axis concerns iterative, step-conditioned retrieval for multi-hop tasks, including the risk of errors accumulating between steps.
Those concerns overlap in implementation even though the authors use separate axes to organize the review. A personalized retrieval policy still needs safety checks; a multi-step pipeline still has a latency budget. The taxonomy gives builders a vocabulary for identifying which tension a method addresses before comparing its reported benefits.
The review's evidence and scope
The authors describe a structured review inspired by PRISMA 2020 principles. Their stated search period covers January 2020 through January 2025. Searches across Google Scholar, Semantic Scholar and arXiv yielded approximately 780 records at the identification stage. That figure is a search-stage count, not a claim that all 780 papers appear in the final synthesis.
Inclusion criteria require a retrieval-and-generation pipeline, a relevant technical or empirical contribution, a RAG evaluation method or work on security and related concerns. Standalone retrievers and language models without integration fall outside those criteria. The authors also describe a targeted citation-snowballing check, using representative papers for each axis and the foundational RAG work as starting points.
The declared search window limits how to read the survey. It supplies an organizing structure for the covered literature, rather than an exhaustive inventory of every release available at publication time.
Retrieval scores do not certify evidence
The paper distinguishes sparse keyword-based retrieval such as BM25 from dense embedding-based matching such as DPR. It also explains a probabilistic formulation that weights generation across retrieved documents.
A useful caveat accompanies that formulation: normalizing scores over the selected top-k documents does not produce a calibrated relevance estimate over the whole corpus. Missing evidence remains missing. The authors also note that document-by-document marginalization is one modeling choice; other systems concatenate context or use reranking and critique scores.
For a team choosing a RAG design, those distinctions prevent a mathematical description from becoming a blanket prescription. The review encourages evaluating the retrieval and generation choices against a named constraint, with source quality and error handling kept in view alongside answer accuracy.
Comments