MATH · OLMOCR-BENCH · OLD SCANS MATH
Math — 90.2% on olmOCR-bench · Old Scans Math.
Old Scans Math measures OCR on equations from scanned mathematical texts — typeset, handwritten, and mixed. Math in non-Latin scripts is especially hard; we score it directly.
Ranked scoreboard.
Models ranked highest-to-lowest. Datalab variants in accent; competitors and prior generations in muted ink.
Rank Model Score
01 Datalab API 90.2%
02 Chandra 2 OSS 89.3%
03 Chandra 1 prior generation 80.3%
+9.0 vs prior generation
olmOCR-bench unit-tests OCR output against known-correct elements in a public dataset. Model versions pinned per run; anyone can reproduce from the olmOCR-bench HuggingFace dataset.
OTHER BENCHMARKS
All benchmarks →Compare on another doc type.
ArXiv Research papers · AI training corpora 90.4% 90.4% on ArXiv, +8.2 vs Chandra 1. Multi-column layout, inline equations, citation graphs — the substrate of most modern AI training data. Tables Financial filings · regulatory PDFs · research data 90.7% Top of the public olmOCR-bench leaderboard. +2.7 vs Chandra 1. Handles colspan, rowspan, nested headers, merged cells.
START
Printed, handwritten, mixed-script equations.
Run a math PDF on Chandra — free tier, no credit card.