Skip to content
← Back to publicationsInformation retrieval

Fusing rankings with incomparable scores

Cormack, Clarke & Büttcher, “Reciprocal Rank Fusion Outperforms Condorcet and Individual Rank Learning Methods”, SIGIR 2009

Combining several retrieval systems requires making their scores comparable. A BM25 score is unbounded and corpus-dependent; a cosine similarity is bounded but concentrated in a narrow range whose distribution varies by model. Any linear combination therefore requires prior normalisation and a weighting, both of which must be estimated on labelled data, which generally does not exist for a client corpus and would have to be re-estimated at every deployment.

Cormack, Clarke and Büttcher proposed in 2009 to sidestep the problem by discarding scores altogether: reciprocal rank fusion uses only each document's rank in each ranking, and sums the reciprocals of those ranks offset by a constant. A document ranked well by several systems obtains a high fusion score even without topping any of them; a document alone at the head of a single system carries less weight. The authors show that this rule, with no learned parameters, outperforms supervised learning-to-rank methods as well as Condorcet fusion.

DeepView fuses the lexical, dense and graph-derived rankings by this method. Our implementation adds per-source weighting to the original formula: each contribution is multiplied by a weight belonging to the retriever it came from, which allows the lexical leg to count for more on a heavily normative corpus than on a set of reports written in free-form language. Those weights and the offset constant are deployment settings, shared across the instance, not values derived automatically from each corpus. The graph contribution further assumes the graph is provisioned; without it, fusion runs over the two remaining rankings.

Going into detail

Architectural choices, the departures we make from the literature, and interim evaluation results are discussed directly with the team that builds the system.