Cross-encoder re-ranking
Nogueira & Cho, “Passage Re-ranking with BERT”, 2019 · arXiv:1901.04085
Dense retrieval must encode each passage independently of the query to allow offline indexing. That constraint carries a precision cost: a passage's representation is computed without knowing the question, and therefore without being able to foreground the aspect of the passage that bears on it. The resulting ranking has good recall, the correct answer nearly always falling within the top hundred results, but mediocre precision at the head of the list.
Nogueira and Cho established in 2019 the effectiveness of a two-stage architecture. A cross-encoder processes the concatenation of query and passage in a single pass, which lets attention operate between the terms of each and yields a markedly more reliable relevance score. Since that computation is quadratic and not indexable, it cannot be applied to the whole corpus; applied to the few dozen candidates surfaced by the first stage, its cost becomes acceptable. The authors report substantial gains on MS MARCO.
DeepView applies re-ranking of this kind to the candidates produced by rank fusion, and passes only the top-scoring ones to the generator. The motivation is twofold: context precision directly conditions generation quality, and a short, dense context lowers the probability that the model grounds its answer in an off-topic passage.
Going into detail
Architectural choices, the departures we make from the literature, and interim evaluation results are discussed directly with the team that builds the system.