67.7k+ open-source stars, hardened in production.
Marker, Surya, and Chandra are ours, and the same engines run inside the commercial platform.
Leading research labs rely on Datalab to accurately turn messy PDFs and other documents into clean, reliable training data. The same open-source models (Marker, Surya, Chandra), hardened, customized, and supported by the team that builds them.








Multi-column, equations, citations, and figures preserved
Try in playground →D · 02 Conference proceedingsLayout-aware, multi-column, multi-figure
Try in playground →D · 03 Lab notebooksMixed handwriting, tables, code blocks
Try in playground →D · 04 PatentsClaims, drawings, and structured metadata
Try in playground →D · 05 Web crawlsHTML to clean training markdown, 90+ languages, at scale
Try in playground →D · 06 Internal RFCsOn-prem parsing of unpublished research
Try in playground →Marker, Surya, and Chandra are ours, and the same engines run inside the commercial platform.
85.9% on olmOCR-bench and 89.9% on tables, with the methodology and competitors published in full.
Our managed batch system can scale to throughput-optimized workloads, already in production and being used by large AI labs processing >100M pages daily.
We train our own models, and can adapt or make new ones tailored to your document processing needs.
Start with an API key — nothing to host or operate.
Sign up →Run in-region, with no egress to US infrastructure.
Sign up →Runs inside your own AWS, GCP, or Azure account.
Talk to sales →Fully offline, the same model weights, dedicated support.
Talk to sales →Free tier with up to $20 in credits per month — no card required.