# How Nevada County Historical Archive made 200K+ pages of Gold Rush history searchable with Datalab

> How the Nevada County Historical Archive uses Datalab's Chandra model to transcribe 200K+ pages of handwritten Gold Rush-era records 150X faster, making a century of California history searchable.

- Canonical: https://www.datalab.to/blog/datalab-nevada-county-historical-archive-case-study
- Published: 2026-06-05
- Authors: Datalab Team

## At a glance

- 150X faster handwritten document transcription (from 3 weeks to 2 hours for a 200-page court case)
- 200K pages indexed to date
- ~95-98% accuracy on modern handwriting
- ~90% accuracy on complex historical handwriting

## Introduction to Nevada County Historical Archive

Dom Lindars is a retired software engineer turned archivist, historian, and author. After his time in Silicon Valley, he began volunteering at the Nevada County Historical Archive, a California-based nonprofit. Their mission is to preserve and celebrate the history of Nevada County, which was central to the Gold Rush in the 1800s. The historical archive has a large set of old records dating back to the 1850s, including mining property logs, land deeds, newspapers, and photos. The Archive can be seen at: [https://archive.nevadacountyhistory.org/](https://archive.nevadacountyhistory.org/)

The Archive is based in Nevada County, California and works with 15 different regional partners, including in Sierra County. Given the imminent threat of wildfires in the area, the Archive has been prioritizing the digitization of historical records to ensure its preservation.

## The challenge: Digitizing handwritten documents accurately and efficiently

In 2025, Dom built a custom Archive to digitize over 800K historical documents. The goal was twofold: preservation and discovery. He solved preservation by scanning all 800K pages, creating digital backups that survive in case of a wildfire disaster. Making them searchable was the next step to allow local researchers and people around the world to easily discover information in the documents.

To do so, he built a PostgreSQL highlight search pipeline around Tesseract's word-level boxes. Initially, the Archive worked for printed material, but he experienced two challenges:

**1. Accuracy on complex documents**

Tesseract worked well on printed documents, but was unable to accurately process handwritten documents. The remaining 150-200K documents of 19th century cursive court cases, property records, mining claims, and records came out unintelligible. Attempting manual transcription of this volume of documents would be unrealistic.

**2. Contextual understanding**

While it produced workable results on printed documents, the output consisted of unstructured word streams, losing formats like columns and tables. On documents like forms and ledgers, the lack of contextual understanding made it difficult to fully understand the output without re-reading the original image.

Searchable digitization is key to how the Nevada County Historical Archive preserves their documents. Nevada County welcomes thousands of tourists and members of the local community every year who are eager to learn about the history of gold mining in the area, going back to the Gold Rush. Many of them come looking for old relatives or information on historical properties, driving hours to the county clerk only to be faced with piles of dusty books and scrawled cursive.

<p style="text-align: center; margin-bottom: 0.5rem; color: #6b7280;"><em>Example of an 1850s handwritten document</em></p>

<img src="/images/blog/datalab-nevada-county-historical-archive-case-study/clinch-family-document.jpg" alt="An 1850s handwritten document from the Searls Historical Library Clinch-Tremoureux Family Collection" style="display: block; margin: 0 auto 2rem; max-height: 680px; width: auto; max-width: 100%; border-radius: 8px;" />

## The solution: Datalab as the handwriting intelligence layer

Dom found Datalab while searching for newer OCR models on Hugging Face. He integrated Datalab's Chandra model via API to convert scans of handwritten documents into structured HTML output with full layout understanding. He tested Chandra on a handful of court case pages, found the transcription to be very high quality, and started routing handwritten material through it.

> "I'm an expert at reading our 19th-century cursive, but I'm consistently surprised at the correct extraction of even badly handwritten letters. Many of my colleagues under 40 say, 'Dom, that's hieroglyphics.' This is lifechanging." — Dom Lindars, Archivist & Technology Lead, Nevada County Historical Archive

Their pipeline runs ingested images through Chandra to produce structured HTML transcripts, which feed PostgreSQL full-text search. The HTML preserves the structural details that are crucial for preserving the context of archival records, like table layouts in assessment ledgers. This allows the records to remain useful as documents beyond a stream of text.

Why Nevada County Historical Archive chose Datalab over alternatives:

- **Searchable handwriting transcription:** Datalab treats handwriting as a first-class input rather than degraded print. The Archive currently experiences ~90% accuracy on difficult 19th century cursive and 95-98% accuracy on cleaner handwriting.
- **Structure-aware output:** The HTML markup output preserves table structure, dollar amounts including cents, fractions, checkboxes, and underlines. For historical documents, structure is an essential part of the data, providing researchers the context they need to understand the documents.
- **Efficiency to address the backlog:** The Archive has processed 200K pages to date - a quarter of its 800K-page collection - in about a month. A 200-300 page handwritten court case that previously took Dom three weeks to transcribe by hand now finishes in 1-2 hours.

<p style="text-align: center; margin-bottom: 0.5rem; color: #6b7280;"><em>Handwritten court case with Datalab-transcribed output</em></p>

<img src="/images/blog/datalab-nevada-county-historical-archive-case-study/people-v-katzenstein-parse.svg" alt="Animation of a handwritten 1858 court case being transcribed by Datalab's Chandra model into structured output" style="display: block; margin: 0 auto 2rem; width: 100%; max-width: 100%; border-radius: 8px;" />

## Impact

Beyond the time savings achieved, making these records searchable has enabled Dom and the Archive to unlock important pieces of history. The Archive receives many requests for volunteers to look up relatives or property records. For most people, it's the difference between driving to the county office and looking through hundreds of pages of old cursive, or typing in key words in the Archive's search bar. Every week, the Archive connects folks with long-lost ancestors and relatives based on the information in the scanned documents.

For example, Aaron A. Sargent, the former California senator and Nevada City resident who wrote the Pacific Railroad Act and drafted the language of the 19th Amendment, was known to be in the Nevada County records. With the records now being searchable, his name is findable in property sale records that previously sat in 80,000 pages of unsearchable microfilm.

## Looking forward

While the Archive has indexed about 200K pages to date, there are an additional 600K pages to process. With Datalab's upcoming release of word-level bounding boxes, the Archive will also extend the same keyword-highlight search to the 200K handwritten pages that now sit alongside them, which will make search highlighting more precise. Currently, searches return a paragraph-level hit, while this update will particularly hit home for users who are looking for exact names or terms on a page, like a relative or an address.
