BENCHMARKS Chandra leads open-source OCR. The Datalab API ships a tuned variant. ← All benchmarks
ARXIV · OLMOCR-BENCH · ARXIV

ArXiv — 90.4% on olmOCR-bench · ArXiv.

ArXiv preprints are dense, multi-column, equation-heavy, and the substrate for most modern AI training corpora. olmOCR-bench checks output against known-correct elements at the page level.

SCOREBOARD · olmOCR-bench · ArXiv

Ranked scoreboard.

Models ranked highest-to-lowest. Datalab variants in accent; competitors and prior generations in muted ink.

Rank Model Score vs scale
01 Datalab API 90.4%
02 Chandra 2 OSS 90.2%
03 Chandra 1 prior generation 82.2%
Dataset · olmOCR-bench · ArXiv Last run · 2026-03-18

+8.0 vs prior generation

WHY THIS NUMBER IS CREDIBLE

olmOCR-bench unit-tests OCR output against known-correct elements in a public dataset. Model versions pinned per run; anyone can reproduce from the olmOCR-bench HuggingFace dataset.

START

Two-column layouts, inline equations, citations.

Run a research paper on Chandra — free tier, no credit card.