ARXIV · OLMOCR-BENCH · ARXIV
ArXiv — 90.4% on olmOCR-bench · ArXiv.
ArXiv preprints are dense, multi-column, equation-heavy, and the substrate for most modern AI training corpora. olmOCR-bench checks output against known-correct elements at the page level.
Ranked scoreboard.
Models ranked highest-to-lowest. Datalab variants in accent; competitors and prior generations in muted ink.
Rank Model Score
01 Datalab API 90.4%
02 Chandra 2 OSS 90.2%
03 Chandra 1 prior generation 82.2%
+8.0 vs prior generation
olmOCR-bench unit-tests OCR output against known-correct elements in a public dataset. Model versions pinned per run; anyone can reproduce from the olmOCR-bench HuggingFace dataset.
OTHER BENCHMARKS
All benchmarks →Compare on another doc type.
Math Academic papers · textbooks · scanned worksheets 90.2% 90.2% on Old Scans Math, +9.9 vs Chandra 1. Printed, handwritten, and mixed-script equations in a single pass. Multilingual 43 languages · training corpora · global ops 80.4% 80.4% across 43 languages. Gemini 2.5 Flash sits at 67.6%, GPT-5 Mini at 60.5%. Bradley-Terry pairwise, Gemini-as-judge, reproducible.
START
Two-column layouts, inline equations, citations.
Run a research paper on Chandra — free tier, no credit card.