Representing meaning in a vector space
Reimers & Gurevych, “Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks”, EMNLP 2019 · arXiv:1908.10084
Lexical matching operates on surface forms. Two passages describing the same concept with disjoint vocabularies (an incident report mentioning an “emergency stop”, a procedure describing “bringing the installation to a safe state”) score zero overlap. This vocabulary mismatch is a long-identified problem in information retrieval, and it worsens on corpora written by many authors over long periods.
Reimers and Gurevych proposed in 2019 a siamese architecture producing sentence representations that are directly comparable. BERT, in its original formulation, must jointly encode both texts to compare them, which makes the cost quadratic in the number of passages and rules out any prior indexing. By training two weight-sharing encoders on a contrastive objective, the authors obtain independent representations, computable offline across the whole corpus and comparable by dot product or cosine similarity. Comparison becomes a nearest-neighbour search, sub-linear with an approximate index.
DeepView indexes the corpus with a model from this family, in a multilingual variant that projects different languages into a shared space: a query in French matches an English passage with no intermediate translation. These models are openly licensed and run on your hardware; both indexing and inference stay inside your environment, with no third-party service involved.
Going into detail
Architectural choices, the departures we make from the literature, and interim evaluation results are discussed directly with the team that builds the system.