TABLES · OLMOCR-BENCH
Tables — 90.7% on olmOCR-bench.
olmOCR-bench unit-tests OCR output against known-correct table cells in real-world documents — financial filings, scientific papers, regulatory reports. We focus on the structurally tricky cases competitors flub.
Ranked scoreboard.
Models ranked highest-to-lowest. Datalab variants in accent; competitors and prior generations in muted ink.
Rank Model Score
01 Datalab API 90.7%
02 Chandra 2 OSS 89.9%
03 Chandra 1 prior generation 88.0%
+1.9 vs prior generation
olmOCR-bench unit-tests OCR output against known-correct elements in a public dataset. Model versions pinned per run; anyone can reproduce from the olmOCR-bench HuggingFace dataset.
OTHER BENCHMARKS
All benchmarks →Compare on another doc type.
Multilingual 43 languages · training corpora · global ops 80.4% 80.4% across 43 languages. Gemini 2.5 Flash sits at 67.6%, GPT-5 Mini at 60.5%. Bradley-Terry pairwise, Gemini-as-judge, reproducible. Math Academic papers · textbooks · scanned worksheets 90.2% 90.2% on Old Scans Math, +9.9 vs Chandra 1. Printed, handwritten, and mixed-script equations in a single pass.
START
Nested headers, merged cells, colspans.
Run your hardest table on Chandra — free tier, no credit card.