BLOG · PRODUCT UPDATES

By Paul Scemama, Hunter Heidenreich, Vikas Paruchuri 9 mins

OmniParseBench: a clear, auditable and explorable benchmark for OCR

Which parser is best on documents like yours? 16,288 tagged pass/fail tests let you check.

Choosing a document parser usually starts with a leaderboard: one number per parser. It is difficult, however, to go a level deeper and understand where and how parsers fail or succeed.

We wanted a benchmark that is:

  • clear: every test is a yes/no question about a page, every score is the share of tests passed, and every score comes with the number of tests and documents behind it, so a thin slice says it is thin and shows where to add tests next;
  • auditable: every test records how its answer was established, and the scorer is open code with no model judging the outputs; and
  • explorable: every test is tagged by what it checks, so any slice of the benchmark, like handwritten tables or Thai text, gets its own score.

Today we’re releasing OmniParseBench: 16,288 pass/fail tests on 2,937 pages from 2,343 documents, in 92 languages, with every test’s answer traced to its source. The dataset and the code are open.

olmOCR-bench showed that pass/fail unit tests are a simple and flexible proxy for parsing quality that anyone can check. But after using it since it came out, we’ve noticed it has gotten saturated, and what now separates the top systems is largely output convention and overfitting rather than reading. Its tests also each sit in one bucket, set by where its page came from, so a weakness on mixed content, like a handwritten table or an equation in a table cell, is averaged into whichever bucket its page landed in.

Another popular benchmark, ParseBench, splits parsing into five capabilities and gives each its own specific metric. Its scores don’t share a unit, so they can’t be compared or combined cleanly, and its tags describe whole pages rather than content. The multiple ways to judge created fairness and reliability issues with ParseBench scores.

Our benchmark

We follow olmOCR-bench in that we use unit tests in order to provide a proxy for real-world parsing performance. A test is a yes/no question about one or more pieces of content on a page. There are three higher-level concepts that describe a test:

  1. Its derivation: how it was verified.
  2. Its type: how it evaluates the content.
  3. Its tags: what kind of content it’s evaluating.

The design

The headline averages three families: text (present, order, repeat tests), tables (table_cell) and layout (layout_kind). A test is one yes/no claim about one page; its args are content at one place on the page; tags on each arg (structure, rendering, role, script, language) select tests for any query and never weigh the headline.
fig. 1 — The benchmark's design: headline, families, tests, args and tags.

The headline is the mean of three family scores: text, tables and layout. Families are just a natural grouping of test types:

test typechecksfamily
presentthis text, number or equation appearstext
ordertext A comes before text Btext
repeatthis line appears exactly N timestext
table_cellthis cell sits under these headings and beside these neighborstables
layout_kindthe block holding this line is a heading, text or tablelayout

A test checks one or more args, each of which is the content at one place on the page. Each arg is tagged:

  • structure: table, math, form and multi_column;
  • rendering: handwriting, tiny, rotated and degraded;
  • role: heading, caption, footnote, list, code, figure and header_footer;
  • script: ISO 15924 codes (Latn, Hani, Deva, …); and
  • language: ISO 639-3 codes (eng, fra, tha, …).

Tags never determine the headline metric or if a test passes or not. Tags are created after a test is made and are for the purpose of querying tests you care about.

Data

The dataset consists of 16,288 tests on 2,937 pages from 2,343 source documents from 22 suites.

Tests, pages and documents per tag. Structure: table 6,767 tests on 1,000 pages from 875 documents; multi_column 3,517; math 3,316; form 1,028. Rendering: handwriting 923, tiny 832, rotated 824, degraded 655. Role: list 861, heading 749, header_footer 227, footnote 104, caption 81, figure 69, code 16 tests on 6 pages from 6 documents. Language: english 10,089 tests on 1,930 pages from 1,612 documents; multilingual (any language but English) 1,931 tests on 529 pages from 321 documents.
fig. 2 — Tests per tag, and the pages and documents they come from.

Tags combine. Here we show a sample of counts for pairs of tags, noting that a test can have an arbitrary number of tags.

A grid of every pair of the 15 tags and multilingual (any language but English), each cell the number of documents with a test carrying both. multi_column with table has 164 documents; math with table 12; handwriting with table 11; 22 pairs have no test, among them math with form, math with handwriting and math with multilingual.
fig. 3 — Documents with a test carrying both tags, for every pair of tags.
All 92 languages by documents and tests. English has 1,612 documents and 10,089 tests; then Mandarin Chinese 50 documents, French 36, Spanish 21, Russian 15, Italian 9, German 8 and Indonesian 7; 32 languages have one document.
fig. 4 — Every language, by the documents and tests it has.

Results

Every score here is based on the dataset in Hugging Face.

parserheadlinetexttableslayout
Gemini 3.8 Flash*92.991.395.292.2
Datalab accurate92.790.891.895.6
Datalab balanced91.690.089.295.5
Reducto90.885.491.295.8
LlamaParse90.488.292.690.4
Mistral OCR89.582.891.993.8
Claude Sonnet 5.5*88.793.195.477.5
GPT-5.6 Sol87.885.491.686.5
Extend83.364.192.593.4
Azure83.074.681.792.8
Tesseractnone17.60.0unsupported

* Used in creating and verifying at least some tests.

Exploring results

Pick test types and tags below, and each parser is scored on the tests you picked the way the headline is: the mean of its pass rates in each family. The scores come from the leaderboard’s runs.

Loading the tests…

To pick tests and aggregate scores your own way on your own runs, see the repository’s guide to querying results.

Accuracy against latency and cost

Headline score against median seconds per page for each parser, with the frontier of the best headline at each latency.
fig. 5 — Headline against median seconds per page, 8 pages in flight.

Latency is the time a client waits for one page: upload, queue, processing and polling. We measured it on 200 pages sampled from the benchmark, sending 8 at a time to each parser, with every parser in its own run.

Headline score against dollars per 1,000 pages for each parser, with the frontier of the best headline at each cost.
fig. 6 — Headline against price per 1,000 pages.

Cost is price per 1,000 pages. For a vendor that bills in credits, it is the credits its responses state times its price per credit. For a general-purpose model we calculate from the tokens it was billed.

Comparing parsers on some tags

A grid of parsers by tags: each cell is a parser's score on the tests with that tag minus the median of the reported parsers on those tests.
fig. 7 — Tag score minus the median of the reported parsers.
START

Get started in minutes.

Free tier. No credit card. SOC 2 Type II.