BENCHMARKS Chandra leads open-source OCR. The Datalab API ships a tuned variant. ← All benchmarks
MULTILINGUAL · INTERNAL · TOP 43 LANGUAGES

Multilingual — 80.4% on Internal · top 43 languages.

There is no good public multilingual OCR benchmark, so we built one. It tests tables, math, ordering, layout, and text accuracy across the top 43 world languages — intentionally hard, to leave headroom. We also publish results on a 90-language long tail.

SCOREBOARD · Internal · top 43 languages

Ranked scoreboard.

Models ranked highest-to-lowest. Datalab variants in accent; competitors and prior generations in muted ink.

Rank Model Score vs scale
01 Datalab API 80.4%
02 Chandra 2 OSS 77.8%
03 Chandra 1 prior generation 69.4%
04 Gemini 2.5 Flash 67.6%
05 GPT-5 Mini 60.5%
Dataset · Internal · top 43 languages Last run · 2026-03-18

+12.8 vs Gemini 2.5 Flash

WHY THIS NUMBER IS CREDIBLE

We built this benchmark — no good public equivalent exists. Pairwise Bradley-Terry with Gemini-as-judge, randomized to reduce position bias, converted to ELO. 95% confidence intervals per row.

START

Gemini gets 67.6%. We get 80.4%.

Run a non-English PDF on Chandra — free tier, no credit card.