PROCESSORS · EVAL

You can't tell if a parse got worse by reading it.

At real volume you can't eyeball conversion quality. Eval scores your output against a rubric you define, so a regression shows up as a falling number — not a downstream surprise.

EVAL · WORKED EXAMPLE

Grade each parse against the answers you already know.

The right answers are in the document — net income, the tables, the order the sections read in. A rubric pins those expectations, and each new run is scored against them, so you can see quality move rather than guess at it.

Alphabet · Q1 2026 10-Q — public domain scan 1 HDR2 HDR3 H14 TBL5 TEXT6 FTR
CHANDRA · PARSE Alphabet · Q1 2026 10-Q · 6 blocks
CONFIDENCE 0.0%
  • 1 Page Header Page Header 0.967
  • 2 Page Header Page Header 0.986
  • 3 Section Header Alphabet Inc. CONSOLIDATED STATEMENTS OF COMPREHENSIVE INCOME (in millions; unaudited) 0.986
  • 4 Table Three Months Ended March 31, 2025 2026 Net income $ 34,540 $ 62,578 Other comprehensive income (loss): Change… 0.960
  • 5 Text See accompanying notes. 0.975
  • 1 more regions detected
chandra · v1.4.2 · 1,287 ms · markdown · json · html
CAPABILITIES

What Eval gives you.

V · 01 RUBRICS

You describe what a good parse looks like

Write a rubric item for each thing that matters — field accuracy, table fidelity, reading order, footnote binding. Each item is graded against the output on a 0–5 scale.

V · 02 SCORING

Every output gets a number, not a gut check

Run the rubric over a set of conversion outputs and get a score per item plus an overall. The documents that pulled the score down are listed, so you can see exactly what broke.

V · 03 REFERENCE CORPUS

Pull the documents that trip you up

Build a corpus from your own production traffic — the messy scans, the dense tables, the formats that fail quietly. That set becomes the bar you measure every parse against.

V · 04 MONITORING

Re-run it on a schedule to catch drift

Document distributions shift and pipelines change over time. Run the same rubric periodically and the score history shows whether your output is holding, improving, or slipping.

API · RUBRIC

Define your scored rubrics.

Each item carries its own 0–5 grade, and they roll up into one overall number you can track from run to run.

# Author a rubric in the UI, then run it over your
# conversion outputs. Each item is graded on a 0–5 scale.

#   Rubric: "Financial Filings · v3"
#     - Tables: column alignment + numeric accuracy   →  4.6 / 5
#     - Net income field extracted correctly          →  5.0 / 5
#     - Reading order across multi-column pages       →  4.2 / 5
#     - Footnote binding                              →  3.8 / 5
#     ─────────────────────────────────────────────
#     Overall:                                            4.4 / 5

# Re-run on a schedule. The score history tells you
# whether parse quality is holding or starting to slip.
RUBRIC · 0–5 PER ITEM · TRACKED OVER TIME
DEPLOYMENT

Run it wherever your data has to live.

  • Managed cloud

    Start with an API key — nothing to host or operate.

    Sign up →
  • EU data residency

    Run in-region, with no egress to US infrastructure.

    Sign up →
  • Your VPC

    Runs inside your own AWS, GCP, or Azure account.

    Talk to sales →
  • On-prem & air-gapped

    Fully offline, the same model weights, dedicated support.

    Talk to sales →
SOC 2 Type II · BAA available View our trust center →
START

Stop guessing whether your parses got worse.

Free tier, no credit card. Author a rubric and score your first run.