Lexical weighting and rare terms
Robertson & Zaragoza, “The Probabilistic Relevance Framework: BM25 and Beyond”, Foundations and Trends in Information Retrieval, 2009
Dense retrieval matches texts on semantic proximity, which makes it robust to variation in phrasing but poorly suited to low-frequency, high-specificity terms: a standard reference, a batch number, a contract identifier. Projected into a low-dimensional continuous space, such terms end up near strings that are graphically close but unrelated. Yet these are precisely the queries where an approximate match has no value at all.
BM25 formalises lexical relevance within the probabilistic framework surveyed by Robertson and Zaragoza. The scoring function combines three components: term frequency within the document, saturated by a parameter that bounds the gain from repeated occurrences; inverse document frequency, which weights a term by its rarity across the corpus; and a document-length normalisation that corrects earlier models' bias toward long documents. No parameter is learned from labelled data: the score is deterministic and inspectable.
DeepView runs BM25 in parallel with dense retrieval, not as a substitute for it. A note on what that buys and what it does not: documents are broken into terms, with no notion of order or case, so a reference such as “PRO-042” counts as two separate terms rather than one string. This is not exact string search. What inverse document frequency does guarantee is that a rare term weighs far more than a common one: a document carrying the right reference rises, where semantic proximity alone would return documents close in subject but carrying a different reference. The two rankings are then fused, by the method described in the entry on reciprocal rank fusion.
Going into detail
Architectural choices, the departures we make from the literature, and interim evaluation results are discussed directly with the team that builds the system.