Assessing retrieval before generating
Yan et al., “Corrective Retrieval Augmented Generation”, 2024 · arXiv:2401.15884
In a standard RAG architecture, generation is triggered whatever passages come back. Answer quality is therefore bounded by retrieval quality, and an upstream failure does not produce a visible downstream failure: the generator, required to answer from an off-topic context, produces an equally off-topic answer phrased with the same confidence. The error is all the more costly for being undetectable by the user.
Yan and co-authors introduced in 2024 a lightweight evaluator applied to retrieved passages before any generation, classifying retrieval into three regimes: correct, incorrect, or ambiguous. Each regime maps to a distinct action: use the passages after fine-grained filtering, trigger an alternative retrieval, or combine both. The mechanism is corrective in the strict sense: it acts on the generator's input rather than its output, and its cost stays low since the evaluator is far lighter than the generator.
DeepView assesses the surviving passages before writing, and re-runs the search when too few remain to answer from: the question is rewritten from a different angle, with broadened vocabulary. Our criterion is cruder than the paper's, which sorts retrieval into three regimes from a score: we count the passages that survive filtering and compare that against a minimum. The number of retries is bounded, which guarantees termination. If no pass brings back enough to answer from, the system says so explicitly rather than writing from passages it has itself judged insufficient. A second departure from the original work: the authors provide for a fallback to web search, which we do not implement, the corpus having to remain the sole scope of any answer.
Going into detail
Architectural choices, the departures we make from the literature, and interim evaluation results are discussed directly with the team that builds the system.