Separating memory from reasoning
Lewis et al., “Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks”, NeurIPS 2020 · arXiv:2005.11401
A language model encodes what it saw during training in its parameters, in a diffuse and non-consultable form. This parametric memory has three structural limits: it is frozen at the collection date, it cannot be edited without retraining, and it supports no attribution, since the model cannot indicate where what it asserts comes from. Queried about a corpus it has never seen, it does not detect the absence: its training objective pushes it toward the most likely continuation, which yields a well-formed and ungrounded answer.
Lewis and co-authors formalised in 2020 an architecture that separates the two functions: a non-parametric memory, consisting of a searchable index over the corpus, and a generative model that conditions its output on passages retrieved at query time. The model no longer has to memorise the corpus, only to write from a supplied context. The authors report gains on knowledge-intensive tasks, and above all a property absent from purely parametric approaches: the passages that conditioned the generation are identifiable.
This is DeepView's architecture, not a layer added on top: generation is systematically conditioned on retrieved passages, never on the model's parametric memory. Two consequences follow. Attribution is structural rather than reconstructed after the fact, since the cited passages are exactly those handed to the generator. And the corpus never has to be learned: the generation model stays generic and interchangeable, with domain knowledge supplied by the index built inside your environment.
Going into detail
Architectural choices, the departures we make from the literature, and interim evaluation results are discussed directly with the team that builds the system.