Reconstructing a PDF's logical structure
Auer et al., “Docling Technical Report”, IBM Research, 2024 · arXiv:2408.09869
PDF is a presentation format: it specifies glyphs and their coordinates on a page, without encoding logical structure. Nothing in it distinguishes a heading from body text, delimits a table, or establishes reading order. Naive extraction concatenates glyphs in stream order, which interleaves the columns of a multi-column layout, linearises a table into a sequence of cells stripped of its headers, and loses section hierarchy. The degradation is silent: the extracted text remains readable while having lost the relations that carried its meaning.
The IBM Research team described in 2024, with Docling, a pipeline that treats layout analysis as a vision problem. An object detection model segments the page into typed regions (heading, paragraph, table, figure, header) whose reading order is then reconstructed; a second, specialised model recovers table structure, including merged cells and hierarchical headers. An optical character recognition engine takes over on pages with no text layer. The output is a structured representation, not a character stream.
This is the first step of indexing in DeepView, and it conditions everything that follows: a row-column relation lost at this stage is unrecoverable, whatever the quality of later steps. Passage segmentation follows this structure rather than a fixed length, which avoids splitting a table or separating a heading from the section it introduces. The whole pipeline runs on your hardware, scanned documents included, with no outbound network access.
Going into detail
Architectural choices, the departures we make from the literature, and interim evaluation results are discussed directly with the team that builds the system.