<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>Datalab Blog</title>
    <link>https://www.datalab.to/blog</link>
    <atom:link href="https://www.datalab.to/blog/rss.xml" rel="self" type="application/rss+xml" />
    <description>Product releases, benchmarks, and case studies from Datalab — document intelligence APIs for conversion, OCR, and structured extraction.</description>
    <language>en</language>
    <item>
      <title>Datalab Document Agent Beta</title>
      <link>https://www.datalab.to/blog/document-agent-beta</link>
      <guid isPermaLink="true">https://www.datalab.to/blog/document-agent-beta</guid>
      <pubDate>Mon, 24 Aug 2026 00:00:00 GMT</pubDate>
      <description>Datalab Document Agent is now in closed beta: specify what you want, provide sample documents, iterate on what correct means, and our agent will devise a custom suite of verifiers and tools built on our state of the art models and pipelines to produce the right document workflow for your needs.</description>
    </item>
    <item>
      <title>Form filling v2 - significantly more accurate</title>
      <link>https://www.datalab.to/blog/form-filling-on-the-document-agent</link>
      <guid isPermaLink="true">https://www.datalab.to/blog/form-filling-on-the-document-agent</guid>
      <pubDate>Mon, 24 Aug 2026 00:00:00 GMT</pubDate>
      <description>Form filling now runs on our document agent: geometry is measured rather than guessed, and every fill is checked against the page it produced. Building the benchmark to prove it also turned up three defects in the version we had been shipping, all of them silent.</description>
    </item>
    <item>
      <title>Segmentation is now $0.50 per 1,000 pages</title>
      <link>https://www.datalab.to/blog/segment-pricing</link>
      <guid isPermaLink="true">https://www.datalab.to/blog/segment-pricing</guid>
      <pubDate>Wed, 19 Aug 2026 00:00:00 GMT</pubDate>
      <description>We cut the price of document segmentation by 92% — and for the most common use, by 95%.</description>
    </item>
    <item>
      <title>Tagged, accessible PDFs in minutes</title>
      <link>https://www.datalab.to/blog/wcag-accessible-pdfs</link>
      <guid isPermaLink="true">https://www.datalab.to/blog/wcag-accessible-pdfs</guid>
      <pubDate>Wed, 19 Aug 2026 00:00:00 GMT</pubDate>
      <description>A new processor that turns a PDF — including a scan that is nothing but pixels — into a tagged PDF built for assistive technology, checked against a battery of automated tests and a real screen reader.</description>
    </item>
    <item>
      <title>Datalab leads another competitor's extraction benchmark</title>
      <link>https://www.datalab.to/blog/extractbench-scoring-bug</link>
      <guid isPermaLink="true">https://www.datalab.to/blog/extractbench-scoring-bug</guid>
      <pubDate>Thu, 13 Aug 2026 00:00:00 GMT</pubDate>
      <description>The benchmark shipped with major scoring bugs. After fixing them, we score 93.6% on the benchmark, leading all other contenders. We encourage you to run your own evals, and not just trust vendor benchmarks.</description>
    </item>
    <item>
      <title>From scanned papers to publisher-ready JATS XML</title>
      <link>https://www.datalab.to/blog/pdf-to-jats-xml</link>
      <guid isPermaLink="true">https://www.datalab.to/blog/pdf-to-jats-xml</guid>
      <pubDate>Thu, 13 Aug 2026 00:00:00 GMT</pubDate>
      <description>A new processor that turns raw scientific paper PDFs into DTD-validated JATS 1.2 XML, with every run carrying a conformance verdict you can audit.</description>
    </item>
    <item>
      <title>ApparelWerks turns handwritten measurements into a system where every inch counts</title>
      <link>https://www.datalab.to/blog/apparelwerks-datalab-casestudy</link>
      <guid isPermaLink="true">https://www.datalab.to/blog/apparelwerks-datalab-casestudy</guid>
      <pubDate>Mon, 10 Aug 2026 00:00:00 GMT</pubDate>
      <description>How a made-to-measure clothing studio used Datalab to build a production system where every inch counts.</description>
    </item>
    <item>
      <title>Improved localization on our customers' hardest documents</title>
      <link>https://www.datalab.to/blog/word-bboxes-degradation-density-rotation</link>
      <guid isPermaLink="true">https://www.datalab.to/blog/word-bboxes-degradation-density-rotation</guid>
      <pubDate>Tue, 04 Aug 2026 00:00:00 GMT</pubDate>
      <description>Bounding boxes for bad scans, dense pages, and sideways text.</description>
    </item>
    <item>
      <title>Improving High Accuracy Mode</title>
      <link>https://www.datalab.to/blog/high-accuracy-update</link>
      <guid isPermaLink="true">https://www.datalab.to/blog/high-accuracy-update</guid>
      <pubDate>Fri, 24 Jul 2026 00:00:00 GMT</pubDate>
      <description>High accuracy mode is now more accurate while making fewer changes to the page, powered by new verification models trained on-distribution with Chandra.</description>
    </item>
    <item>
      <title>Marker 2: faster, CPU-ready, and more accurate</title>
      <link>https://www.datalab.to/blog/marker-2</link>
      <guid isPermaLink="true">https://www.datalab.to/blog/marker-2</guid>
      <pubDate>Mon, 20 Jul 2026 00:00:00 GMT</pubDate>
      <description>Marker 2 is a rewrite of our open-source PDF-to-markdown converter. It runs on CPU, picks a speed/accuracy mode automatically, and beats comparable pipeline OCR systems on both accuracy and throughput on olmOCR-bench.</description>
    </item>
    <item>
      <title>When a citation needs to survive a courtroom</title>
      <link>https://www.datalab.to/blog/kin-datalab-casestudy</link>
      <guid isPermaLink="true">https://www.datalab.to/blog/kin-datalab-casestudy</guid>
      <pubDate>Fri, 17 Jul 2026 00:00:00 GMT</pubDate>
      <description>How Kin rebuilt the way adjusters cite insurance policies, achieving compliance-grade accuracy with Chandra and custom processors.</description>
    </item>
    <item>
      <title>We lead an external extraction benchmark (but run your own evals)</title>
      <link>https://www.datalab.to/blog/trusting-vendor-benchmarks</link>
      <guid isPermaLink="true">https://www.datalab.to/blog/trusting-vendor-benchmarks</guid>
      <pubDate>Tue, 14 Jul 2026 00:00:00 GMT</pubDate>
      <description>A competitor commissioned a long-table extraction benchmark, and scored Datalab at 34% recall. We rebuilt our extraction and now have the top recall (99.1%) and lead on precision too.  Still, you should take all vendor benchmarks with a grain of salt and run your own evals.</description>
    </item>
    <item>
      <title>Learnboost fixes the layer beneath the AI and lifts student retention</title>
      <link>https://www.datalab.to/blog/learnboost-datalab-casestudy</link>
      <guid isPermaLink="true">https://www.datalab.to/blog/learnboost-datalab-casestudy</guid>
      <pubDate>Wed, 08 Jul 2026 00:00:00 GMT</pubDate>
      <description>How an AI-powered learning platform boosted retention by replacing its document conversion with Marker.</description>
    </item>
    <item>
      <title>Tell us how you want your documents parsed</title>
      <link>https://www.datalab.to/blog/custom-processors</link>
      <guid isPermaLink="true">https://www.datalab.to/blog/custom-processors</guid>
      <pubDate>Tue, 07 Jul 2026 00:00:00 GMT</pubDate>
      <description>Custom Processors turn plain-language instructions and a few examples into a processor that parses your documents exactly the way you need.</description>
    </item>
    <item>
      <title>How we achieve a near-zero OCR hallucination rate</title>
      <link>https://www.datalab.to/blog/anti-hallucination</link>
      <guid isPermaLink="true">https://www.datalab.to/blog/anti-hallucination</guid>
      <pubDate>Tue, 30 Jun 2026 00:00:00 GMT</pubDate>
      <description>Hallucinations are a fact of how LLMs work. For mission-critical documents, even one is too many. Here's the defense-in-depth system we use to drive them toward zero.</description>
    </item>
    <item>
      <title>A box and a confidence score for every word</title>
      <link>https://www.datalab.to/blog/word-bounding-boxes-and-confidence</link>
      <guid isPermaLink="true">https://www.datalab.to/blog/word-bounding-boxes-and-confidence</guid>
      <pubDate>Tue, 23 Jun 2026 00:00:00 GMT</pubDate>
      <description>Making OCR output auditable.</description>
    </item>
    <item>
      <title>Free to start, pay as you go</title>
      <link>https://www.datalab.to/blog/free-to-start-pay-as-you-go</link>
      <guid isPermaLink="true">https://www.datalab.to/blog/free-to-start-pay-as-you-go</guid>
      <pubDate>Thu, 18 Jun 2026 00:00:00 GMT</pubDate>
      <description>Datalab now has a free tier and pay-as-you-go pricing. Start parsing for free with a monthly usage allowance, then pay only for the pages you actually process — no subscription, no minimum, no plan to pick.</description>
    </item>
    <item>
      <title>Introducing lift: open-weights structured extraction</title>
      <link>https://www.datalab.to/blog/introducing-lift</link>
      <guid isPermaLink="true">https://www.datalab.to/blog/introducing-lift</guid>
      <pubDate>Thu, 18 Jun 2026 00:00:00 GMT</pubDate>
      <description>lift is a 9B open-weights vision model that extracts structured JSON from PDFs and images by passing a schema. It scores 90.2% field accuracy on our 225-document extraction benchmark - the best small self-hostable model we've tested.</description>
    </item>
    <item>
      <title>Chandra 2.1: Improved Multilingual and Table Accuracy</title>
      <link>https://www.datalab.to/blog/chandra-2.1-release</link>
      <guid isPermaLink="true">https://www.datalab.to/blog/chandra-2.1-release</guid>
      <pubDate>Wed, 17 Jun 2026 00:00:00 GMT</pubDate>
      <description>Announcing Chandra 2.1, a smaller, faster model that improves on multilingual and table accuracy.</description>
    </item>
    <item>
      <title>Turbo Mode: Structured Extraction At 12s/doc</title>
      <link>https://www.datalab.to/blog/turbo-extraction</link>
      <guid isPermaLink="true">https://www.datalab.to/blog/turbo-extraction</guid>
      <pubDate>Thu, 11 Jun 2026 00:00:00 GMT</pubDate>
      <description>Turbo mode for the Datalab extraction API: structured JSON from any document at a median of ~12 seconds per document.</description>
    </item>
    <item>
      <title>How Nevada County Historical Archive made 200K+ pages of Gold Rush history searchable with Datalab</title>
      <link>https://www.datalab.to/blog/datalab-nevada-county-historical-archive-case-study</link>
      <guid isPermaLink="true">https://www.datalab.to/blog/datalab-nevada-county-historical-archive-case-study</guid>
      <pubDate>Fri, 05 Jun 2026 00:00:00 GMT</pubDate>
      <description>How the Nevada County Historical Archive uses Datalab's Chandra model to transcribe 200K+ pages of handwritten Gold Rush-era records 150X faster, making a century of California history searchable.</description>
    </item>
    <item>
      <title>Introducing Balanced Mode for Structured Extraction</title>
      <link>https://www.datalab.to/blog/introducing-balanced-extraction-mode</link>
      <guid isPermaLink="true">https://www.datalab.to/blog/introducing-balanced-extraction-mode</guid>
      <pubDate>Thu, 04 Jun 2026 00:00:00 GMT</pubDate>
      <description>Balanced is our new default extraction mode: higher accuracy, with reasoning and independent verification baked into every response. It knows when a value isn't in the document instead of making one up.</description>
    </item>
    <item>
      <title>EU Data Residency now available on Datalab</title>
      <link>https://www.datalab.to/blog/eu-data-residency</link>
      <guid isPermaLink="true">https://www.datalab.to/blog/eu-data-residency</guid>
      <pubDate>Fri, 29 May 2026 00:00:00 GMT</pubDate>
      <description>Datalab now offers data residency controls, starting with a new processing location in the Netherlands. Keep your document content and processing entirely within the EU, using the same API you already use.</description>
    </item>
    <item>
      <title>Announcing Surya OCR 2: small, accurate, multilingual</title>
      <link>https://www.datalab.to/blog/surya-2</link>
      <guid isPermaLink="true">https://www.datalab.to/blog/surya-2</guid>
      <pubDate>Wed, 27 May 2026 00:00:00 GMT</pubDate>
      <description>Surya OCR 2 is a 650M-parameter open-source OCR model that scores 83.3% on olmOCR-bench, hits 87.2% on a 91-language multilingual eval, and runs on CPU, GPU, and MPS.</description>
    </item>
    <item>
      <title>Managed Batch Processing is now Live</title>
      <link>https://www.datalab.to/blog/managed-batch-processing-is-now-live</link>
      <guid isPermaLink="true">https://www.datalab.to/blog/managed-batch-processing-is-now-live</guid>
      <pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate>
      <description>Introducing Datalab's Managed Batch Processing — give us a bucket, we handle the rest. Process millions of pages without managing GPUs, scaling infrastructure, or orchestrating workloads.</description>
    </item>
    <item>
      <title>How Radical AI Accelerates Materials Science Discovery with Datalab</title>
      <link>https://www.datalab.to/blog/datalab-radical-ai-case-study</link>
      <guid isPermaLink="true">https://www.datalab.to/blog/datalab-radical-ai-case-study</guid>
      <pubDate>Thu, 26 Mar 2026 00:00:00 GMT</pubDate>
      <description>Radical AI ($55M Seed+) ingests millions of pages of dense academic literature through Datalab to feed the autonomous-discovery systems shrinking materials R&amp;D from decades to weeks.</description>
    </item>
    <item>
      <title>Announcing Chandra OCR 2: 90+ Languages, Top Benchmarks</title>
      <link>https://www.datalab.to/blog/chandra-2</link>
      <guid isPermaLink="true">https://www.datalab.to/blog/chandra-2</guid>
      <pubDate>Wed, 18 Mar 2026 00:00:00 GMT</pubDate>
      <description>Chandra 2 is a 4B parameter OCR model with state of the art benchmarks, layout blocks with bounding boxes, and structured output for diagrams and charts.</description>
    </item>
    <item>
      <title>Chandra 1.5: Better Tables, Chemistry Support, and Faster Processing</title>
      <link>https://www.datalab.to/blog/chandra-1-5</link>
      <guid isPermaLink="true">https://www.datalab.to/blog/chandra-1-5</guid>
      <pubDate>Thu, 22 Jan 2026 00:00:00 GMT</pubDate>
      <description>Announcing Chandra 1.5 with dramatically improved table extraction, chemistry support with SMILES output, diagram rendering, and significant latency improvements.</description>
    </item>
    <item>
      <title>How Rely trusts Datalab to process massive data rooms in hours, not days</title>
      <link>https://www.datalab.to/blog/datalab-rely-case-study</link>
      <guid isPermaLink="true">https://www.datalab.to/blog/datalab-rely-case-study</guid>
      <pubDate>Fri, 16 Jan 2026 00:00:00 GMT</pubDate>
      <description>Rely audits commercial real-estate transactions worth millions in misvalued assets. Running Datalab on-prem at 12.5 pages/sec lets them ingest 500,000-page data rooms before deals close — without resident documents ever leaving their network.</description>
    </item>
    <item>
      <title>How Cofactr (YC W'22) Automates Complex Logistics Workflows to Support Critical Hardware Supply Chains</title>
      <link>https://www.datalab.to/blog/cofactr-datalab-casestudy</link>
      <guid isPermaLink="true">https://www.datalab.to/blog/cofactr-datalab-casestudy</guid>
      <pubDate>Tue, 06 Jan 2026 00:00:00 GMT</pubDate>
      <description>Cofactr (YC W'22) serves regulated hardware manufacturers — Amazon Robotics among them. Datalab parses inconsistent BOMs, quotes, and spec sheets across 5,000+ suppliers with the fidelity their compliance and traceability work demands.</description>
    </item>
    <item>
      <title>Extracting Hyperlinks from PDFs</title>
      <link>https://www.datalab.to/blog/extracting-hyperlinks-from-pdfs</link>
      <guid isPermaLink="true">https://www.datalab.to/blog/extracting-hyperlinks-from-pdfs</guid>
      <pubDate>Fri, 19 Dec 2025 00:00:00 GMT</pubDate>
      <description>Learn how to extract hyperlinks from PDFs using the Datalab API, preserving both link text and destinations.</description>
    </item>
    <item>
      <title>Introducing Form Filling: Automatically Fill PDF Forms with AI</title>
      <link>https://www.datalab.to/blog/automatically-fill-pdf-forms-with-ai</link>
      <guid isPermaLink="true">https://www.datalab.to/blog/automatically-fill-pdf-forms-with-ai</guid>
      <pubDate>Wed, 17 Dec 2025 00:00:00 GMT</pubDate>
      <description>Automatically fill PDF forms using AI with our new form filling feature.</description>
    </item>
    <item>
      <title>Introducing: Forge Evals</title>
      <link>https://www.datalab.to/blog/introducing-forge-evals</link>
      <guid isPermaLink="true">https://www.datalab.to/blog/introducing-forge-evals</guid>
      <pubDate>Wed, 10 Dec 2025 00:00:00 GMT</pubDate>
      <description>A new tool to evaluate how various settings and modes impact document parsing quality across PDFs, DOCX files, and spreadsheets.</description>
    </item>
    <item>
      <title>Datalab Benchmarks + Evals: View Scores &amp; Compare Output On Your Documents</title>
      <link>https://www.datalab.to/blog/datalab-benchmarks-evals</link>
      <guid isPermaLink="true">https://www.datalab.to/blog/datalab-benchmarks-evals</guid>
      <pubDate>Tue, 09 Dec 2025 00:00:00 GMT</pubDate>
      <description>View scores and compare model output across a sample of diverse pages, or get evals on your own documents.</description>
    </item>
    <item>
      <title>Launch Week - Day 5: Faster Tracked Changes Outputs</title>
      <link>https://www.datalab.to/blog/launch-week-faster-tracked-changes-outputs</link>
      <guid isPermaLink="true">https://www.datalab.to/blog/launch-week-faster-tracked-changes-outputs</guid>
      <pubDate>Fri, 05 Dec 2025 00:00:00 GMT</pubDate>
      <description>Performance improvements for Tracked Changes output parsing, plus native spreadsheet support in Forge.</description>
    </item>
    <item>
      <title>Launch Week - Day 4: Spreadsheet Parsing</title>
      <link>https://www.datalab.to/blog/launch-week-spreadsheet-parsing</link>
      <guid isPermaLink="true">https://www.datalab.to/blog/launch-week-spreadsheet-parsing</guid>
      <pubDate>Thu, 04 Dec 2025 00:00:00 GMT</pubDate>
      <description>Native spreadsheet support in the Datalab API for accurate table segmentation and parsing.</description>
    </item>
    <item>
      <title>Launch Week - Day 3: Introducing Agni: Solving Multi-Page Section Hierarchy in OCR</title>
      <link>https://www.datalab.to/blog/launch-week-introducing-agni-solving-multi-page-section-hierarchy-in-ocr</link>
      <guid isPermaLink="true">https://www.datalab.to/blog/launch-week-introducing-agni-solving-multi-page-section-hierarchy-in-ocr</guid>
      <pubDate>Wed, 03 Dec 2025 00:00:00 GMT</pubDate>
      <description>Introducing Agni, a new model that maintains consistent section hierarchy across multi-page documents.</description>
    </item>
    <item>
      <title>Launch Week - Day 2: Chandra is Faster (Again)</title>
      <link>https://www.datalab.to/blog/launch-week-chandra-is-faster-again</link>
      <guid isPermaLink="true">https://www.datalab.to/blog/launch-week-chandra-is-faster-again</guid>
      <pubDate>Tue, 02 Dec 2025 00:00:00 GMT</pubDate>
      <description>Announcing Chandra Small, a latency-optimized model that achieves 2-3x faster speeds with minimal performance degradation.</description>
    </item>
    <item>
      <title>Launch Week - Day 1: Chandra 1.1</title>
      <link>https://www.datalab.to/blog/launch-week-chandra-1-1</link>
      <guid isPermaLink="true">https://www.datalab.to/blog/launch-week-chandra-1-1</guid>
      <pubDate>Mon, 01 Dec 2025 00:00:00 GMT</pubDate>
      <description>Announcing Chandra 1.1 with improved layout, math, tables, and multilingual performance.</description>
    </item>
    <item>
      <title>Speeding Up Chandra</title>
      <link>https://www.datalab.to/blog/speeding-up-chandra</link>
      <guid isPermaLink="true">https://www.datalab.to/blog/speeding-up-chandra</guid>
      <pubDate>Mon, 17 Nov 2025 00:00:00 GMT</pubDate>
      <description>How we achieved 3x faster Chandra performance using Eagle3 speculative decoding without sacrificing accuracy.</description>
    </item>
    <item>
      <title>Saturating the olmOCR Benchmark</title>
      <link>https://www.datalab.to/blog/saturating-the-olmocr-benchmark</link>
      <guid isPermaLink="true">https://www.datalab.to/blog/saturating-the-olmocr-benchmark</guid>
      <pubDate>Thu, 13 Nov 2025 00:00:00 GMT</pubDate>
      <description>Our analysis reveals fundamental limitations in the olmOCR benchmark and why traditional edit-distance benchmarks provide diminishing returns for OCR evaluation.</description>
    </item>
    <item>
      <title>Extract Tracked Changes Metadata from Word Documents into Markdown &amp; HTML</title>
      <link>https://www.datalab.to/blog/extract-tracked-changes-metadata</link>
      <guid isPermaLink="true">https://www.datalab.to/blog/extract-tracked-changes-metadata</guid>
      <pubDate>Wed, 12 Nov 2025 00:00:00 GMT</pubDate>
      <description>New feature enabling extraction of tracked changes metadata from Word documents for contract review workflows.</description>
    </item>
    <item>
      <title>Grounded Intelligence: How High-Fidelity OCR Drives Accurate Structured Extraction</title>
      <link>https://www.datalab.to/blog/high-fidelity-ocr-drives-accurate-structured-extraction</link>
      <guid isPermaLink="true">https://www.datalab.to/blog/high-fidelity-ocr-drives-accurate-structured-extraction</guid>
      <pubDate>Tue, 04 Nov 2025 00:00:00 GMT</pubDate>
      <description>How superior OCR technology enhances the accuracy of structured data extraction from documents.</description>
    </item>
    <item>
      <title>Introducing our newest model: Chandra</title>
      <link>https://www.datalab.to/blog/introducing-chandra</link>
      <guid isPermaLink="true">https://www.datalab.to/blog/introducing-chandra</guid>
      <pubDate>Thu, 30 Oct 2025 00:00:00 GMT</pubDate>
      <description>Meet Chandra, our latest OCR model that tops independent benchmarks and brings powerful new capabilities to document processing.</description>
    </item>
    <item>
      <title>View Your API Requests in Our Playground</title>
      <link>https://www.datalab.to/blog/view-your-api-requests-in-our-playground</link>
      <guid isPermaLink="true">https://www.datalab.to/blog/view-your-api-requests-in-our-playground</guid>
      <pubDate>Tue, 07 Oct 2025 00:00:00 GMT</pubDate>
      <description>Visualize your API requests directly within the Playground to audit and debug your document processing.</description>
    </item>
    <item>
      <title>Launch Week Day 5: Playground Examples</title>
      <link>https://www.datalab.to/blog/launch-week-playground-examples</link>
      <guid isPermaLink="true">https://www.datalab.to/blog/launch-week-playground-examples</guid>
      <pubDate>Tue, 30 Sep 2025 00:00:00 GMT</pubDate>
      <description>We've added practical examples to the Playground that you can explore directly.</description>
    </item>
    <item>
      <title>Launch Week Day 4: High Accuracy Mode</title>
      <link>https://www.datalab.to/blog/launch-week-high-accuracy-mode</link>
      <guid isPermaLink="true">https://www.datalab.to/blog/launch-week-high-accuracy-mode</guid>
      <pubDate>Fri, 26 Sep 2025 00:00:00 GMT</pubDate>
      <description>Introducing High Accuracy Mode, which combines our proprietary models with frontier LLMs to handle challenging documents.</description>
    </item>
    <item>
      <title>Launch Week Day 3: Layout Model Updates</title>
      <link>https://www.datalab.to/blog/launch-week-layout-model</link>
      <guid isPermaLink="true">https://www.datalab.to/blog/launch-week-layout-model</guid>
      <pubDate>Thu, 25 Sep 2025 00:00:00 GMT</pubDate>
      <description>Significant upgrades to our OCR and Layout models, achieving state-of-the-art performance on olmOCR-bench.</description>
    </item>
    <item>
      <title>Launch Week Day 2: Document Segmentation</title>
      <link>https://www.datalab.to/blog/launch-week-document-segmentation</link>
      <guid isPermaLink="true">https://www.datalab.to/blog/launch-week-document-segmentation</guid>
      <pubDate>Wed, 24 Sep 2025 00:00:00 GMT</pubDate>
      <description>New Document Segmentation feature that automatically detects page boundaries within multi-document PDF files.</description>
    </item>
    <item>
      <title>Launch Week Day 1: Datalab's New Playground</title>
      <link>https://www.datalab.to/blog/launch-week-new-playground</link>
      <guid isPermaLink="true">https://www.datalab.to/blog/launch-week-new-playground</guid>
      <pubDate>Tue, 23 Sep 2025 00:00:00 GMT</pubDate>
      <description>Unveiling our updated Playground with parsing, refinement, segmentation, and extraction capabilities.</description>
    </item>
    <item>
      <title>Cracking Math OCR: How We're Unlocking High-Quality Data for Reasoning Models</title>
      <link>https://www.datalab.to/blog/cracking-math-ocr</link>
      <guid isPermaLink="true">https://www.datalab.to/blog/cracking-math-ocr</guid>
      <pubDate>Fri, 19 Sep 2025 00:00:00 GMT</pubDate>
      <description>How high-quality mathematical text extraction enables training of reasoning-capable language models.</description>
    </item>
    <item>
      <title>Structured Extraction with Datalab API, and Handling Long Documents</title>
      <link>https://www.datalab.to/blog/structured-extraction-with-datalab-api-and-handling-long-documents</link>
      <guid isPermaLink="true">https://www.datalab.to/blog/structured-extraction-with-datalab-api-and-handling-long-documents</guid>
      <pubDate>Mon, 08 Sep 2025 00:00:00 GMT</pubDate>
      <description>How to use Datalab's Marker API for structured extraction from PDFs using JSON schemas, plus strategies for processing lengthy documents.</description>
    </item>
    <item>
      <title>Citation Needed - Auditable Structured Extraction</title>
      <link>https://www.datalab.to/blog/structured-extraction-citations</link>
      <guid isPermaLink="true">https://www.datalab.to/blog/structured-extraction-citations</guid>
      <pubDate>Tue, 19 Aug 2025 00:00:00 GMT</pubDate>
      <description>Structured Extraction now includes citation support, enabling users to verify information sources within documents.</description>
    </item>
    <item>
      <title>Parse PDFs *Just the Way You Want*</title>
      <link>https://www.datalab.to/blog/parse-pdfs-just-the-way-you-want</link>
      <guid isPermaLink="true">https://www.datalab.to/blog/parse-pdfs-just-the-way-you-want</guid>
      <pubDate>Thu, 14 Aug 2025 00:00:00 GMT</pubDate>
      <description>Introducing Marker Prompt API and Forge Parse for customizing PDF parsing outputs through prompts.</description>
    </item>
    <item>
      <title>Introducing Forge Extract: Turn PDFs into Structured JSON</title>
      <link>https://www.datalab.to/blog/releases-forge-extract</link>
      <guid isPermaLink="true">https://www.datalab.to/blog/releases-forge-extract</guid>
      <pubDate>Thu, 07 Aug 2025 00:00:00 GMT</pubDate>
      <description>Extract structured data from PDFs using custom schemas with our new Forge Extract feature.</description>
    </item>
    <item>
      <title>Purchaser.ai: Six-figure cost savings and 50% higher retention with Datalab</title>
      <link>https://www.datalab.to/blog/datalab-purchaser-case-study</link>
      <guid isPermaLink="true">https://www.datalab.to/blog/datalab-purchaser-case-study</guid>
      <pubDate>Wed, 16 Jul 2025 00:00:00 GMT</pubDate>
      <description>Manufacturing procurement runs on messy quotes and thousand-row BOMs. Datalab handles the inconsistency so Purchaser.ai's team ships faster purchasing decisions — and keeps the customers who'd otherwise churn on bad parses.</description>
    </item>
    <item>
      <title>The Datalab SDK: Transform Document Processing from Hours to Minutes</title>
      <link>https://www.datalab.to/blog/releases-sdk</link>
      <guid isPermaLink="true">https://www.datalab.to/blog/releases-sdk</guid>
      <pubDate>Wed, 09 Jul 2025 00:00:00 GMT</pubDate>
      <description>Announcing the official Datalab SDK for Python with simplified APIs and powerful features.</description>
    </item>
    <item>
      <title>RevisionDojo (YC F24): Achieving 10x cost reduction and seamless scalability with Datalab</title>
      <link>https://www.datalab.to/blog/datalab-revisiondojo-case-study</link>
      <guid isPermaLink="true">https://www.datalab.to/blog/datalab-revisiondojo-case-study</guid>
      <pubDate>Tue, 01 Jul 2025 00:00:00 GMT</pubDate>
      <description>RevisionDojo's edtech platform serves 310K+ students who need study materials parsed on demand. Retiring their self-hosted screenshot pipeline for Datalab cut processing cost 10x and brought average parse time to ~10 seconds.</description>
    </item>
    <item>
      <title>Gamma: Unlocking a 40% Conversion Boost with reliable and accurate PDF conversion</title>
      <link>https://www.datalab.to/blog/datalab-gamma-case-study</link>
      <guid isPermaLink="true">https://www.datalab.to/blog/datalab-gamma-case-study</guid>
      <pubDate>Tue, 17 Jun 2025 00:00:00 GMT</pubDate>
      <description>Gamma's 50M users turn PDFs into working presentations. Switching parsers to Datalab lifted PDF-import conversion 40% — over a million PDFs a week now flow through the engine without dropping layout or table fidelity.</description>
    </item>
    <item>
      <title>Building better document intelligence for an AI-first world</title>
      <link>https://www.datalab.to/blog/building-better-document-intelligence-for-an-ai-first-world</link>
      <guid isPermaLink="true">https://www.datalab.to/blog/building-better-document-intelligence-for-an-ai-first-world</guid>
      <pubDate>Thu, 12 Jun 2025 00:00:00 GMT</pubDate>
      <description>Why highly accurate document intelligence matters more than ever in an AI-first world.</description>
    </item>
  </channel>
</rss>