Skip to content
← Back to publicationsReliability and verification

Verifying an answer's groundedness in its sources

Asai et al., “Self-RAG: Learning to Retrieve, Generate and Critique through Self-Reflection”, ICLR 2024 · arXiv:2310.11511

Conditioning generation on retrieved passages reduces the hallucination rate without eliminating it: nothing in the generation objective requires each assertion produced to be entailed by the context. The model may extrapolate beyond what a passage establishes, or ignore a relevant passage outright. These two failure modes are distinct from retrieval failures, and invisible to the user, since the answer remains well-formed and accompanied by plausible citations.

Asai and co-authors proposed at ICLR 2024 to train the model to emit, alongside the text, reflection tokens that make its own judgement explicit: whether retrieval is needed for this segment, whether the retrieved passages are relevant, whether the generated segment is supported by them, and whether the answer is useful to the question asked. These judgements, produced during decoding, allow segments to be selected or rejected and make the system's behaviour controllable at inference time. The point that matters for us is the separation of two criteria the literature often conflates: factual support by the sources, and relevance to the question.

DeepView adopts that separation without adopting the implementation. We do not train a model to emit reflection tokens, which would require fine-tuning per deployment: the two criteria are evaluated after generation, by a separate check over the answer and the passages that conditioned it. What those checks then do depends on the effort level in use. At the standard level the answer is streamed as it is written and the two verdicts accompany it for information: they flag a poorly supported answer, they do not withhold it. At the higher levels the answer is held back while it is assessed, and a failure re-runs retrieval from a different angle before anything is shown. This is a latency-versus-verification trade-off, not a settled property of the system.

Going into detail

Architectural choices, the departures we make from the literature, and interim evaluation results are discussed directly with the team that builds the system.