Skip to content

Research

The work DeepView
is built on.

DeepView is a Retrieval-Augmented Generation (RAG) system. Each of its components implements a method published and evaluated by the research community: lexical and dense retrieval, rank fusion, re-ranking, knowledge graph extraction, verification of answer groundedness.

We only draw on work from the field's leading conferences and teams, and only where it holds up against real documents. Each publication has its own entry: the problem it formalises, the method it proposes, and what our implementation retains from it, including the departures we make from the original work.

Document analysis

Śmigielski et al., “Chunking Methods on Retrieval-Augmented Generation: Effectiveness Evaluation Against Computational Cost and Limitations”, Wrocław University of Science and Technology, 2026 · arXiv:2606.00881

Splitting documents into passages looks like a technical setting with no consequences. It is in fact one of the choices that weighs most on what a system can retrieve.

In DeepView: the choice of chunking strategy

Read the article →

Information retrieval

Lewis et al., “Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks”, NeurIPS 2020 · arXiv:2005.11401

A language model's knowledge is frozen in its parameters at training time. Queried outside that distribution, it generates a plausible answer rather than a refusal.

In DeepView: the overall system architecture

Read the article →

Robertson & Zaragoza, “The Probabilistic Relevance Framework: BM25 and Beyond”, Foundations and Trends in Information Retrieval, 2009

Dense representations project rare terms into a continuous space where an identifier loses the specificity that made it discriminating. Lexical weighting remains necessary.

In DeepView: the lexical component of retrieval

Read the article →

Cross-encoder re-ranking

Reliability and verification

Nogueira & Cho, “Passage Re-ranking with BERT”, 2019 · arXiv:1901.04085

Independent encoders are indexable but imprecise: query and passage are vectorised separately, with no interaction between their terms.

In DeepView: re-ranking of candidates

Read the article →

Cormack, Clarke & Büttcher, “Reciprocal Rank Fusion Outperforms Condorcet and Individual Rank Learning Methods”, SIGIR 2009

A BM25 score and a cosine similarity share neither scale nor distribution. Combining them linearly assumes a normalisation that nothing justifies.

In DeepView: rank fusion

Read the article →

Knowledge graph

Edge et al., “From Local to Global: A Graph RAG Approach to Query-Focused Summarization”, Microsoft Research, 2024 · arXiv:2404.16130

Passage retrieval treats each unit independently. A query whose answer is distributed across dozens of documents exceeds what that paradigm can assemble.

In DeepView: graph extraction and traversal

Read the article →

Reliability and verification

Attribution and how we choose

This work belongs to its authors and has no connection to Alister AI: we claim neither authorship nor endorsement, and none of these authors contributed to DeepView. We select a publication on two criteria: how solid its evaluation is, and whether it holds up on real documents rather than on academic benchmarks alone. A gain that does not survive contact with corporate documentation is of no interest to us, however prestigious the venue. That is also why we sometimes depart from the original method where our deployment constraints call for different choices: those departures are flagged in the entries concerned.

Evaluation

Internal evaluation protocol

We evaluate what the knowledge graph adds over a purely vector RAG architecture, on a question set stratified by complexity. The protocol is described below; the campaign is under way.

  • Baseline systema standard vector RAG architecture (AnythingLLM), with the same corpus and generation model.
  • System under testDeepView: hybrid retrieval (BM25, dense, graph) followed by re-ranking.
  • Question setaround 100 questions over an enterprise document corpus, stratified into 5 complexity levels, with manually established reference answers.
  • Metricsanswer accuracy and completeness, groundedness in the cited sources, hallucination rate, multi-hop resolution, latency and cost per query.

Question stratification

  1. Level 1 — Direct extraction

    the answer appears verbatim in a single passage.

  2. Level 2 — Semantic matching

    the question and the relevant passage share no vocabulary.

  3. Level 3 — Aggregation

    the answer requires combining several passages or documents.

  4. Level 4 — Multi-hop

    the answer requires following a chain of relations between entities.

  5. Level 5 — Discriminating questions

    questions built so that vector similarity alone is insufficient.

Status

The evaluation campaign is under way. We will publish the full study here: detailed protocol, question set, and per-level, per-metric results. Until then, the chart opposite represents our working hypothesis, not measurements.

Going into detail

Architectural choices, the departures we make from the literature, and interim evaluation results are discussed directly with the team that builds the system.