BENCHMARKS · OLMOCR-BENCH + INTERNAL EVAL

Chandra leads open-source OCR. The Datalab API goes further.

Chandra hits 85.9% on olmOCR-bench — state of the art among open OCR models. The Datalab API runs a modified Chandra at 86.7%, with extra accuracy, throughput, and an extraction workflow layered on top.

86.7% overall · olmOCR-bench Datalab API · 2026-03-18
TRUSTED BY
OVERALL · OLMOCR-BENCH

86.7% via API · 85.9% open. #1 among open OCR.

Chandra 2 tops the public olmOCR-bench leaderboard at 85.9% open-weights. The Datalab API ships at 86.7% on the same model, with automatic correction layered in. Run your own PDF and verify.

OPEN OCR
  • olmOCR
  • RolmOCR
  • dots.ocr
  • DeepSeek OCR
FRONTIER · MULTILINGUAL
  • Gemini 2.5 Flash
  • GPT-5 Mini
PRIOR GENERATION
  • Chandra 1
Rank Model Score vs scale
01 Datalab API 86.7%
02 Chandra 2 OSS SOTA among open models 85.9%
03 Chandra 1 prior generation 83.1%
Dataset · olmOCR-bench · all categories Last run · 2026-03-18

Frontier LLMs (Gemini 2.5 Flash, GPT-5 Mini) scored on multilingual only — see the Multilingual tab below.

Try Chandra free → Free tier · no credit card
WHAT YOU PARSE

See your document parsed.

Pick the document type closest to yours. Numbers, named competitors, and live parsed examples for each.

90.7% Datalab API · 89.9% open · #1 among open OCR on olmOCR-bench tables Ranked table + methodology →
Document type not here? Legal contracts, medical records, handwritten forms — we parse those too. Talk to sales →
START

Try our hosted API on your hardest PDF.

Free tier. No credit card. 100M+ pages a day, sustained — capacity isn't the bottleneck.