Est.

Building a Ground Truth Dataset for OCR Evaluation

Defining what "correct" means is harder than checking model output against it.

Staff Writer · · 11 min read
Cover illustration for “Building a Ground Truth Dataset for OCR Evaluation”
Accuracy Measurement · October 6, 2026 · 11 min read · 2,400 words

Building a ground truth dataset for OCR evaluation is not an annotation task you hand off and forget. The intuitive view treats this as simple: gather some documents, get the correct answers written down, compare model output against them, done. That view misses where the actual difficulty lives. The hard part isn't checking a model's output against a known answer. The hard part is deciding, field by field and document by document, what the known answer should be, and that decision is loaded with choices about scope, format, and edge cases that are easy to get wrong in ways nobody notices until the model is already deployed.

A ground truth dataset is a specification of what "correct" means for a given extraction task. That specification can be wrong, incomplete, or narrower than the real world, and none of those problems appear in a validation report. Instead, they appear months later, when a model that scored well in testing starts producing errors nobody flagged during evaluation.

This is why the gap between benchmark accuracy and production accuracy is overwhelmingly a dataset construction problem, not a model problem. Each one is a lever. If the lever is pulled wrong, the resulting dataset will quietly validate a system that fails the moment it meets real documents.

The document population you sample determines which failures you will and won't see

Sampling is the first engineering decision, and it's the one that sets a hard ceiling on what the evaluation can ever detect. If a failure mode doesn't exist in the sampled population, no metric downstream will catch it. Each axis produces its own distinct failure profile, and skipping any one of them means skipping the failures that live there.

Layout is a clear example. When OCR engines are evaluated across FUNSD, SROIE, and Rx-PAD, as DocOCR-Eval shows, the same engine scores well on one domain or language and far worse on another. A dataset built from just one of those three benchmarks would hand you a ranking of OCR tools that looks confident and is systematically wrong for any document population that resembles the other two<sup>1</sup><sup>4</sup>.

Domain specificity works the same way, and historical documents make the point sharply. Reichsanzeiger-GT exists because a ground truth built from modern scans cannot cover what the "Deutscher Reichsanzeiger und Preußischer Staatsanzeiger," a German-language newspaper spanning 1819 to 1945, actually looks like: degraded paper, non-standard typography, multilingual content accumulated across more than a century. A separate ground truth covering those specific conditions is required, because no amount of volume in the wrong population substitutes for coverage of the right one.

Image quality deserves its own line of attention, because OCR systems don't degrade gracefully as input quality drops. They degrade non-linearly, meaning performance can hold steady across a range of quality and then collapse sharply past some threshold. That makes image quality a variable to sample on purpose, not a nuisance to filter out by picking only clean-looking scans. Research on document image quality assessment found that standard perceptual quality metrics, including BRISQUE, NIQE, FRIQUEE, and PIQA, don't correlate reliably with actual OCR accuracy. Choosing "good quality" samples by eyeballing them is not a valid stand-in for deliberately sampling across the quality range a production system will actually face.

The failures that cause the most damage in production rarely come from the easy cases. Skipping that audit means every later step inherits its blind spots.

What the transcription standard must settle before annotation begins

Sampling decides what documents go into the dataset. The next decision is how "correct" gets defined once an annotator sits down with one of those documents in front of them. Most ground truth corruption doesn't come from annotators making mistakes. It comes from under-specification: when the transcription standard leaves a case ambiguous, different annotators resolve it differently, each choice locally reasonable, and the resulting dataset carries noise that looks like signal.

A transcription standard has to settle several things explicitly before annotation starts, not during it.

Formatting rules for superscripts, subscripts, special symbols, and mathematical notation need to be fixed in advance. Skipping this means the dataset inherits the exact failure PureDocBench documents: the unit m³ transcribed as "m³3" in one place and "m3" in another, depending on how a given model or annotator interpreted the symbol. In a technical or scientific document, a correct unit and a silently wrong one are at stake.

Whitespace normalization needs an explicit rule. Pick a convention and apply it uniformly, or the ground truth will penalize or reward formatting choices that have nothing to do with actual extraction quality.

Illegible or genuinely ambiguous characters need a defined protocol: mark them, skip them, or assign a sentinel value. Nobody auditing the final dataset will know that a "correct" answer was actually a guess.

Automated ground truth generation, which aligns an electronic source document to its scanned image, removes some of these ambiguities but doesn't eliminate the need for a standard. It shifts where the standard has to be enforced. Automating the matching process doesn't remove the specification work. It moves that work earlier, into the alignment logic.

For domain-specific documents like invoices, payslips, or closing disclosures, the standard has to go one level deeper and define what counts as the value of a field, not just its surface text. What happens when a field appears more than once on the same page? These are decisions that belong in the standard itself, not edge cases to improvise during annotation.

A complete transcription standard still leaves one question open: did the people doing the annotation actually follow it? That's a separate problem, and it needs its own process.

Annotation quality control as a technical process, not a proofreading pass

A ground truth dataset is only as trustworthy as the process used to verify it, and a pipeline that never measures inter-annotator agreement cannot distinguish annotator noise from genuine ambiguity in the documents themselves. Quality control, done right, is a measurement activity. A human skimming a sample and deciding it "looks right" tells you nothing comparable.

Inter-annotator agreement needs to be computed at the same granularity the model itself will be evaluated against. For OCR text recognition, that means character-level agreement. For structured extraction, that means field-level agreement. Measuring agreement at a coarser level than the actual evaluation hides exactly the disagreements that matter most, because two annotators can agree on the large majority of a document's characters while disagreeing on the one field that actually gets used downstream<sup>2</sup><sup>1</sup>.

Field-level agreement, specifically, exposes three distinct kinds of disagreement that a document-level metric would blur together: whether annotators identified the same field boundaries, whether they recorded the same value strings, and whether they applied the same normalization choices. Those are three separate failure sources, and conflating them into one document-level pass rate means you can't tell which part of your standard needs fixing.

Disagreements between annotators are diagnostic signals, not noise to resolve by majority vote and move on. When annotators disagree systematically on a particular document type or field type, that's evidence the transcription standard is under-specified for that case, and the fix is to revise the standard, not to outvote the disagreement and pretend it didn't happen.

Quality control also has to be structurally independent of annotation. Having the same annotators check each other's work without a defined protocol catches random slips but misses systematic errors, because a systematic error is, by definition, one that looks consistent and reasonable to anyone applying the same flawed assumption.

For fields where errors carry real operational cost, amounts, dates, identification numbers, the minimum standard is a second independent annotation pass followed by adjudication from a domain expert. It is the baseline requirement for any field where a wrong value causes downstream harm, not an upgrade reserved for high-budget projects.

Evaluation metrics chosen to match the cost of the error, not the convenience of the formula

Once the ground truth itself is solid, the next engineering decision is which metric to run against it, and that choice carries its own judgment about which errors matter and how much. A well-built ground truth evaluated with the wrong metric still produces a misleading accuracy number. The question to ask for any given field is simple: given what this value will actually be used for, what does "wrong" mean here? That question should drive metric selection, not the other way around.

Character error rate and word error rate treat every character position as equally important. That makes both metrics well suited to measuring OCR fidelity on continuous text, where no single character carries outsized weight. It makes them poorly suited to structured extraction, where a single wrong digit in an invoice total is categorically worse than a stray punctuation mark in a description field. Treating those two errors as equivalent, which CER and WER do by construction, buries the error that actually costs money under the one that doesn't<sup>2</sup><sup>1</sup>.

Field-level accuracy, whether the extracted value for a defined field matches the ground truth value exactly, is the right metric for structured extraction, because it maps directly onto the operational question that matters: will this value get used correctly by whatever system consumes it next?

The math behind multi-field documents makes this non-negotiable. A document with many fields has a substantially lower probability of being entirely correct than any single field's accuracy rate would suggest, simply because the probability of every field being right at once compounds downward as the field count grows.

Matching algorithms matter just as much as the metric formula sitting on top of them. OmniDocBench's Multi-Granularity Adaptive Matching approach addresses a specific failure in naive string matching: a prediction segmented differently than the ground truth, correct in substance but split across different boundaries, can register as wrong under a rigid match even though the content is right. MGAM adapts granularity on the prediction side while leaving the ground truth untouched, which removes that bias without changing what "correct" means.

Normalized edit distance metrics like ANLS handle partial matches more gracefully than a binary field-match check, but they need calibration before they're trusted. A high ANLS score on a currency amount that's off by a single digit is a wrong number that happens to look similar to the right one, not a partial success in any operational sense.

Table-heavy documents need their table structure evaluated separately from their text, using a metric like TEDS, because a table can pass a character-level check while having collapsed merged cells or transposed columns, and that table is wrong in a way that matters. In financial and technical documents, the answer lives in the structure, not just the characters sitting inside it.

Engineering confidence scores calibrated against ground truth rather than self-reported by the model

A confidence score only earns its keep if it predicts error. A model that assigns high confidence to wrong answers and low confidence to correct ones gives no usable signal for routing documents to human review or for setting an automation threshold, and worse, it actively misleads whoever is relying on that threshold to make decisions.

Calibrating confidence requires the same rigor as building the ground truth itself. To trust a confidence score, measure, across a representative sample of extractions, whether fields scored above a given threshold are actually correct at the rate the score implies. That calibration can't happen on the same documents used to train or tune the model. It requires a held-out evaluation set, drawn from the same population the model will actually see in production, built with the same discipline as the main ground truth dataset.

Calibration isn't a single number that transfers everywhere. A threshold tuned on date fields in clean typeset documents won't hold for amount fields in low-resolution scans, and it won't hold for fields in languages the model encountered less often during training.

A high-confidence output isn't evidence of correctness unless that confidence is grounded in the actual pixels on the page. Confidence, in that case, measures fluency, not correctness.

The operational consequence is direct: any automation threshold set on an uncalibrated confidence score amounts to a guess. Because the errors that slip through a high-confidence gate are, by definition, the high-confidence ones, they're also the errors least likely to get flagged by a human reviewer downstream. Percentile routing approves the least-bad documents in a bad batch. It does not approve the documents that are actually correct.

Silent failures that a correctly constructed evaluation set will catch and a typical benchmark will not

The most damaging failures in production OCR and extraction do not appear in standard benchmark evaluation, because standard benchmarks weren't built to surface them. Catching them requires deliberately including the document types and conditions where they actually occur.

Multi-column reading order errors are one example. Catching it requires the ground truth to encode the expected reading order explicitly, with evaluation that checks order, not just whether the right characters showed up somewhere.

Table structure collapse works the same way. The numbers are all present. Their relationships to each other are gone, and in a financial or technical table, those relationships are the entire point.

Silent omission is harder still to catch, because an evaluation that only measures what was extracted cannot detect what was never extracted.

Symbol and special-character misrecognition occurs constantly in technical content: mathematical notation, unit symbols, and typographic marks like superscripts, subscripts, and arrows get misread in ways that look plausible on the surface. The m³ becoming m3 is one instance. An arrow annotation collapsing into a stray character is another. Both produce downstream errors in any system that treats the value semantically.

OCR post-correction can introduce its own failure class on top of all of this. The HIPE-OCRepair-2026 competition identifies this overcorrection as a recurring challenge, and it's a distinct failure from the original OCR error, sometimes a worse one, because the corrected text reads as confident and clean while being wrong in a new way.

Each of these failure modes has a concrete requirement attached to it: explicit reading-order ground truth, table structure annotation kept separate from text accuracy, completeness accounting that captures what was dropped, symbol-level transcription rules settled in advance, and evaluation that checks correction output against the original rather than assuming cleaner text means more accurate text. A dataset built without these specific requirements in mind won't find these failures, no matter how large it is or how many documents it contains.

Sources

  1. Frontiers
  2. DocOCR-Eval: A Correction-Based Framework for OCR Tool Selection Without Ground Truth
  3. ICDAR 2026 HIPE-OCRepair Competition on LLM-Assisted OCR Post-Correction for Historical Documents
  4. Reichsanzeiger-GT: An OCR ground truth dataset based on the historical newspaper “Deutscher Reichsanzeiger und Preußischer Staatsanzeiger” (German Imperial Gazette and Prussian Official Gazette) (1819–1945)
  5. OCR Error Post-Correction with LLMs in Historical Documents: No Free Lunches

More in Accuracy Measurement