Skip to content
← Back to publicationsDocument analysis

Reading a page and its reading order in one pass

Cui et al., “RT-DocLayout: Real-Time End-to-End Document Layout Analysis with Reading Order in the Wild”, 2026 · arXiv:2606.23344

Document analysis pipelines conventionally chain separate models: region detection, segmentation, then reading-order reconstruction. Each stage inherits the errors of the one before it with no way to correct them: a badly cut region stays badly cut, and reading order is computed over blocks that are already wrong. The problem worsens on what the literature calls documents “in the wild”, meaning real ones: pages scanned at an angle, warped paper, photographs taken off-axis, composite layouts. That is exactly the material a system deployed inside a company receives, where the archives hold documents scanned fifteen years ago.

Cui and co-authors propose in 2026 an architecture that unifies classification, detection, pixel-level segmentation and reading-order prediction into a single model. Treating the four tasks jointly removes error propagation between stages: reading order is predicted from the same representation as segmentation, rather than recomputed afterwards over a frozen output. The model reaches state-of-the-art layout analysis with only 33 million parameters, at 132 frames per second, while the field trends towards far heavier vision-language models.

That trade-off interests us more than the benchmark ranking. A 33-million-parameter model runs on hardware our clients already own, with no dedicated graphics card for indexing, where a large vision-language model would demand either new hardware or a call to an outside service, which our deployment principle rules out. Speed matters just as much: migrating an archive means hundreds of thousands of pages, and a tenfold difference in indexing throughput is the difference between one night and two weeks.

Going into detail

Architectural choices, the departures we make from the literature, and interim evaluation results are discussed directly with the team that builds the system.