# OmniParseBench: a clear, auditable and explorable benchmark for OCR

> Which parser is best on documents like yours? 16,288 tagged pass/fail tests let you check.

- Canonical: https://www.datalab.to/blog/omni-parse-bench
- Published: 2026-10-08
- Authors: Paul Scemama, Hunter Heidenreich, Vikas Paruchuri

Choosing a document parser usually starts with a leaderboard: one number per parser. It is difficult, however, to go a level deeper and understand where and how parsers fail or succeed.

We wanted a benchmark that is:

- clear: every test is a yes/no question about a page, every score is the share of tests passed, and every score comes with the number of tests and documents behind it, so a thin slice says it is thin and shows where to add tests next;
- auditable: every test records how its answer was established, and the scorer is open code with no model judging the outputs; and
- explorable: every test is tagged by what it checks, so any slice of the benchmark, like handwritten tables or Thai text, gets its own score.

Today we're releasing OmniParseBench: 16,288 pass/fail tests on 2,937 pages from 2,343 documents, in 92 languages, with every test's answer traced to its source. The [dataset](https://huggingface.co/datasets/datalab-to/omni_parse_bench) and the [code](https://github.com/datalab-to/omni_parse_bench) are open.

[olmOCR-bench](https://huggingface.co/datasets/allenai/olmOCR-bench) showed that pass/fail unit tests are a simple and flexible proxy for parsing quality that anyone can check. But after using it since it came out, we've noticed it has gotten saturated, and what now separates the top systems is largely output convention and overfitting rather than reading. Its tests also each sit in one bucket, set by where its page came from, so a weakness on mixed content, like a handwritten table or an equation in a table cell, is averaged into whichever bucket its page landed in.

Another popular benchmark, [ParseBench](https://www.parsebench.ai/), splits parsing into five capabilities and gives each its own specific metric. Its scores don't share a unit, so they can't be compared or combined cleanly, and its tags describe whole pages rather than content. The multiple ways to judge created fairness and reliability issues with ParseBench scores.

## Our benchmark

We follow olmOCR-bench in that we use unit tests in order to provide a proxy for real-world parsing performance. A test is a yes/no question about one or more pieces of content on a page. There are three higher-level concepts that describe a test:

1. Its _derivation_: how it was verified.
2. Its _type_: how it evaluates the content.
3. Its _tags_: what kind of content it's evaluating.

## The design

<figure style="margin:32px 0">
  <img src="/images/blog/omni-parse-bench/model.svg" alt="The headline averages three families: text (present, order, repeat tests), tables (table_cell) and layout (layout_kind). A test is one yes/no claim about one page; its args are content at one place on the page; tags on each arg (structure, rendering, role, script, language) select tests for any query and never weigh the headline." style="display:block;width:100%;max-width:760px;margin-inline:auto" />
  <figcaption style="font-family:var(--font-mono);font-size:0.7rem;letter-spacing:0.08em;text-transform:uppercase;color:var(--color-ink-3);margin-top:12px">fig. 1 — The benchmark's design: headline, families, tests, args and tags.</figcaption>
</figure>

The headline is the mean of three family scores: text, tables and layout. Families are just a natural grouping of test types:

| test type     | checks                                                         | family |
| ------------- | -------------------------------------------------------------- | ------ |
| `present`     | this text, number or equation appears                          | text   |
| `order`       | text A comes before text B                                     | text   |
| `repeat`      | this line appears exactly N times                              | text   |
| `table_cell`  | this cell sits under these headings and beside these neighbors | tables |
| `layout_kind` | the block holding this line is a heading, text or table        | layout |

A test checks one or more args, each of which is the content at one place on the page. Each arg is tagged:

- structure: `table`, `math`, `form` and `multi_column`;
- rendering: `handwriting`, `tiny`, `rotated` and `degraded`;
- role: `heading`, `caption`, `footnote`, `list`, `code`, `figure` and `header_footer`;
- script: ISO 15924 codes (`Latn`, `Hani`, `Deva`, ...); and
- language: ISO 639-3 codes (`eng`, `fra`, `tha`, ...).

_Tags never determine the headline metric or if a test passes or not_. Tags are created _after_ a test is made and are for the purpose of querying tests you care about.

## Data

The dataset consists of 16,288 tests on 2,937 pages from 2,343 source documents from 22 suites.

<figure style="margin:32px 0">
  <img src="/images/blog/omni-parse-bench/tag-coverage.svg" alt="Tests, pages and documents per tag. Structure: table 6,767 tests on 1,000 pages from 875 documents; multi_column 3,517; math 3,316; form 1,028. Rendering: handwriting 923, tiny 832, rotated 824, degraded 655. Role: list 861, heading 749, header_footer 227, footnote 104, caption 81, figure 69, code 16 tests on 6 pages from 6 documents. Language: english 10,089 tests on 1,930 pages from 1,612 documents; multilingual (any language but English) 1,931 tests on 529 pages from 321 documents." style="display:block;width:100%;max-width:760px;margin-inline:auto" />
  <figcaption style="font-family:var(--font-mono);font-size:0.7rem;letter-spacing:0.08em;text-transform:uppercase;color:var(--color-ink-3);margin-top:12px">fig. 2 — Tests per tag, and the pages and documents they come from.</figcaption>
</figure>

Tags combine. Here we show a sample of counts for _pairs_ of tags, noting that a test can have an arbitrary number of tags.

<figure style="margin:32px 0">
  <img src="/images/blog/omni-parse-bench/tag-pairs.svg" alt="A grid of every pair of the 15 tags and multilingual (any language but English), each cell the number of documents with a test carrying both. multi_column with table has 164 documents; math with table 12; handwriting with table 11; 22 pairs have no test, among them math with form, math with handwriting and math with multilingual." style="display:block;width:100%;max-width:664px;margin-inline:auto" />
  <figcaption style="font-family:var(--font-mono);font-size:0.7rem;letter-spacing:0.08em;text-transform:uppercase;color:var(--color-ink-3);margin-top:12px">fig. 3 — Documents with a test carrying both tags, for every pair of tags.</figcaption>
</figure>

<figure style="margin:32px 0">
  <img src="/images/blog/omni-parse-bench/languages.svg" alt="All 92 languages by documents and tests. English has 1,612 documents and 10,089 tests; then Mandarin Chinese 50 documents, French 36, Spanish 21, Russian 15, Italian 9, German 8 and Indonesian 7; 32 languages have one document." style="display:block;width:100%;max-width:760px;margin-inline:auto" />
  <figcaption style="font-family:var(--font-mono);font-size:0.7rem;letter-spacing:0.08em;text-transform:uppercase;color:var(--color-ink-3);margin-top:12px">fig. 4 — Every language, by the documents and tests it has.</figcaption>
</figure>

## Results

Every score here is based on the dataset in [Hugging Face](https://huggingface.co/datasets/datalab-to/omni_parse_bench/tree/175cc91a9df1a8d1ee2339ad652d3dce20852911).

| parser             | headline | text | tables |      layout |
| ------------------ | -------: | ---: | -----: | ----------: |
| Gemini 3.8 Flash*  |     92.9 | 91.3 |   95.2 |        92.2 |
| Datalab accurate   |     92.7 | 90.8 |   91.8 |        95.6 |
| Datalab balanced   |     91.6 | 90.0 |   89.2 |        95.5 |
| Reducto            |     90.8 | 85.4 |   91.2 |        95.8 |
| LlamaParse         |     90.4 | 88.2 |   92.6 |        90.4 |
| Mistral OCR        |     89.5 | 82.8 |   91.9 |        93.8 |
| Claude Sonnet 5.5* |     88.7 | 93.1 |   95.4 |        77.5 |
| GPT-5.6 Sol        |     87.8 | 85.4 |   91.6 |        86.5 |
| Extend             |     83.3 | 64.1 |   92.5 |        93.4 |
| Azure              |     83.0 | 74.6 |   81.7 |        92.8 |
| Tesseract          |     none | 17.6 |    0.0 | unsupported |

\* Used in creating and verifying at least some tests.

### Exploring results

Pick test types and tags below, and each parser is scored on the tests you picked the way the headline is: the mean of its pass rates in each family. The scores come from the leaderboard's runs.

To pick tests and aggregate scores your own way on your own runs, see [the repository's guide to querying results](https://github.com/datalab-to/omni_parse_bench/blob/main/docs/API.md#asking-a-question).

### Accuracy against latency and cost

<figure style="margin:32px 0">
  <img src="/images/blog/omni-parse-bench/latency.svg" alt="Headline score against median seconds per page for each parser, with the frontier of the best headline at each latency." style="display:block;width:100%;max-width:760px;margin-inline:auto" />
  <figcaption style="font-family:var(--font-mono);font-size:0.7rem;letter-spacing:0.08em;text-transform:uppercase;color:var(--color-ink-3);margin-top:12px">fig. 5 — Headline against median seconds per page, 8 pages in flight.</figcaption>
</figure>

Latency is the time a client waits for one page: upload, queue, processing and polling. We measured it on 200 pages sampled from the benchmark, sending 8 at a time to each parser, with every parser in its own run.

<figure style="margin:32px 0">
  <img src="/images/blog/omni-parse-bench/cost.svg" alt="Headline score against dollars per 1,000 pages for each parser, with the frontier of the best headline at each cost." style="display:block;width:100%;max-width:760px;margin-inline:auto" />
  <figcaption style="font-family:var(--font-mono);font-size:0.7rem;letter-spacing:0.08em;text-transform:uppercase;color:var(--color-ink-3);margin-top:12px">fig. 6 — Headline against price per 1,000 pages.</figcaption>
</figure>

Cost is price per 1,000 pages. For a vendor that bills in credits, it is the credits its responses state times its price per credit. For a general-purpose model we calculate from the tokens it was billed.

### Comparing parsers on some tags

<figure style="margin:32px 0">
  <img src="/images/blog/omni-parse-bench/weak-spots.svg" alt="A grid of parsers by tags: each cell is a parser's score on the tests with that tag minus the median of the reported parsers on those tests." style="display:block;width:100%;max-width:1074px;margin-inline:auto" />
  <figcaption style="font-family:var(--font-mono);font-size:0.7rem;letter-spacing:0.08em;text-transform:uppercase;color:var(--color-ink-3);margin-top:12px">fig. 7 — Tag score minus the median of the reported parsers.</figcaption>
</figure>
