Est.

Per-Field Accuracy Measurement in Document Extraction

Measuring accuracy by individual fields exposes failures hidden in aggregate scores.

Editor at Large · · 12 min read
Cover illustration for “Per-Field Accuracy Measurement in Document Extraction”
Accuracy Measurement · September 29, 2026 · 12 min read · 2,730 words

A single accuracy number for a document extraction system tells you almost nothing about whether that system will work on your documents. That is the core problem with how most vendors report performance, before getting into the mechanics of fixing it.

Take the cloud-seeding extraction study built on 832 NOAA weather reports, which reported an overall accuracy of 98.38 percent Pharma Document Extraction Benchmark: Tables & Footnotes | IntuitionLabs koreadeep.com extend.ai digitalapplied.com Structured dataset of reported cloud seeding activities in the United States (2000-2025) using an LLM. On its face, that's a system operating close to perfection. But break the number down by field, and Season is 87.94 percent, Agent is 89.95 percent, and Apparatus is 92.96 percent, even as Year and State both hit 100 percent Pharma Document Extraction Benchmark: Tables & Footnotes | IntuitionLabs koreadeep.com extend.ai digitalapplied.com Structured dataset of reported cloud seeding activities in the United States (2000-2025) using an LLM. Averaging those together produces a headline figure that flatters the system while hiding the fact that roughly one in eight Season extractions comes back wrong, a field that could carry real regulatory or operational weight depending on what it feeds downstream.

This isn't a quirk specific to weather reports. The same math applies anywhere a pipeline mixes easy fields with hard ones: a system that nails a binary yes/no field nearly every time can drag its average up high enough to bury a categorical field that's wrong a third of the time. Financial, medical, and legal extraction pipelines run into this constantly, because the fields inside a single document type rarely share a difficulty level, let alone an error tolerance.

Production systems don't consume averages. An accounts payable system posts an invoice using the actual value in amount_due. A loan origination system reads the exact figure in the loan amount field. An insurance adjudication system checks a policy_number against a database record. Each of those is a discrete decision made on a discrete value, and if that value is wrong, the aggregate score offers no protection whatsoever.

There's also an incentive problem baked into how most accuracy claims get produced. Vendors choose their own test sets, and they choose which metric to publish. Without a per-field breakdown, a buyer has no way to audit whether the number reflects performance on documents anything like the ones they'll actually run through the system.

Diagram: How a 98% Headline Hides an 87% Field. Visualizes: Show the gap between an aggregate accuracy figure and its per-field breakdown using real numbers from the cloud-seeding study.

What per-field accuracy measurement means in practice

Per-field measurement compares a single extracted value against a verified, human-produced ground truth for that exact field in that exact document, not document against document. A document can be judged "correct" overall while several of its fields are wrong, if the scoring method rewards partial credit or rounds errors away. Field-by-field comparison doesn't allow that kind of averaging to happen quietly.

Building a ground truth set worth trusting takes discipline. That independence check matters because it validates the ground truth itself, not just the model being measured against it. If two trained annotators can't agree with each other on what the correct value should be, that reflects a quality issue in the ground truth itself.

Sample size isn't arbitrary. Scaling that logic to your own document volume and error tolerance turns the sample size from a guess into a calculated figure.

Annotator agreement also sets a practical ceiling that's easy to forget about. Research on footnote boundary detection found human annotators agreeing with each other at a mean Average Precision of 83 to 91 across document categories Pharma Document Extraction Benchmark: Tables & Footnotes | IntuitionLabs. That band isn't a floor to clear, it's closer to the maximum score any automated system should be expected to hit against that kind of ground truth.

The metric itself needs to match the field type, and this is where a lot of evaluation setups go wrong by applying one yardstick everywhere. Character Error Rate suits OCR-sensitive fields like names, addresses, and codes, where a single dropped digit matters. Word Error Rate fits free-text fields where entire words go missing or get substituted. Tree Edit Distance Similarity handles table cells, where the structure of the table matters as much as the content inside it. Using exact-match scoring on a free-text field, or CER on a table, produces numbers that look precise but measure the wrong thing.

None of this is the same job as running a foundation model through a standardized academic benchmark. OCRBench v2, SROIE, and OmniDocBench, the last of which was accepted at CVPR 2025, are genuinely useful for comparing model capabilities in the abstract extend.ai. But none of them will tell a team whether their system handles signature detection correctly, or preserves table structure, on the actual documents that pipeline processes every day extend.ai. Production evaluation has to run on production documents. There's no substitute. The Donohue et al. study selected 200 randomly sampled records out of 832 to achieve a margin of error of ±4% at a 90% confidence level, a concrete benchmark for practitioners calibrating their own evaluation sample Structured dataset of reported cloud seeding activities in the United States (2000-2025) using an LLM Pharma Document Extraction Benchmark: Tables & Footnotes | IntuitionLabs extend.ai.

The three failure modes that per-field confidence scores must catch

Confidence scores are supposed to work like a contract: a field gets accepted automatically if its score clears a threshold, and the system guarantees that the error rate among accepted fields stays below some target level. The trouble is that a lot of teams implement that contract using a procedure that sounds reasonable and turns out not to hold up.

The common approach, sometimes called the folklore procedure, thresholds a confidence score on a calibration split using an add-one bound extend.ai. It has the shape of rigor. A preprint on valid per-field selective risk control shows that it silently violates the contract on real documents⟧c13⟧ extend.ai. The paper tested this against 13,859 genuine fields drawn from 800 CORD receipts, where field correctness sat at 49.0 percent, and diagnosed three distinct failure modes hiding inside that setup.

The first is document clustering. The threshold ends up looking well-calibrated on paper while failing to deliver that calibration on real documents, because the statistics assumed independence that the data never had.

The second failure mode is score-refit leakage: fitting a learned confidence score and its threshold on the same set of fields. This violates the nominal risk level in the vast majority of splits tested. The system's certified coverage claim is one it cannot actually back up in production.

The third is a tie-mass pathology, where a degenerate score distribution collapses the threshold grid and produces near-zero certified coverage from what amounts to a single signal failure.

On the CORD dataset, only 49.0 percent of asserted fields are correct Valid Per-Field Selective Risk Control for Document Extraction: Three Failure Modes, a Validity Ladder, and When Conditioning Pays. Grounding, meaning whether a field value can be traced back to a specific span in the source document, turned out to be strongly discriminative for spotting errors, but grounding alone doesn't produce a certified guarantee on its own Valid Per-Field Selective Risk Control for Document Extraction: Three Failure Modes, a Validity Ladder, and When Conditioning Pays. A field can point convincingly at the right part of the page and still be wrong.

The practical tier, a straightforward fit/val split protocol, controls expected selective risk on average but lets realized risk exceed the target in roughly half of resplits. The rigorous field-iid PAC tier uses Mondrian Learn-then-Test with exact binomial tails, giving per-group certificates, but it assumes field independence that documents routinely violate.

The practical upshot: confidence scores and accept/review thresholds don't certify themselves. The calibration procedure behind the number determines the stated risk level, and most vendor implementations never disclose which procedure they used. Fields within the same document are correlated, and the design effect of 1.84–2.45 roughly halves the effective calibration size, so the threshold appears well-calibrated but isn't digitalapplied.com Valid Per-Field Selective Risk Control for Document Extraction: Three Failure Modes, a Validity Ladder, and When Conditioning Pays. The rigorous doc-iid PAC tier uses per-document bounds, the only tier whose assumptions actually match how documents are structured, but it yields near-vacuous coverage at current model performance levels.

Where field-level extraction breaks in production documents

Failure modes stop being abstract once you look at where they actually show up on real document types. ParseBench, an April 2026 preprint, identifies four failure modes that break production agentic workflows built on document extraction.

Tables fail on structural fidelity: merged cells, hierarchical headers, and content that continues across a page break. A single shifted header can cause an agent to extract the wrong value entirely, and it fails silently in financial workflows where nothing downstream flags the mismatch. Charts fail differently, because exact data-point extraction with correctly matched labels is hard, and agents often receive a descriptive summary in place of a precise value, then act on an input that was never right. Content faithfulness failures appear as omissions and hallucinations, where dropped or fabricated content occurs when the agent reasons over context that was never in the source document. Semantic formatting failures are the subtlest: strikethrough, superscript, subscript, bold, and hyperlinks all carry legal and financial meaning, and dropping them produces output that looks structurally correct while being semantically wrong.

PureDocBench, a 2026 benchmark, names three more failure modes tied directly to field-level consequences. E1 is header metadata omission, where the entire product header block goes missing from reconstruction, leaving downstream users unable to identify which component a datasheet even describes. E2 is a reading-order error, where tables survive intact but land in the wrong section, breaking the logical grouping engineers depend on. E3 is a technical symbol recognition error in which engineering symbols get silently mutated, for example nH rendered as mH, a 10⁶× unit error, a safety-critical misinterpretation that passes every structural check while being factually wrong.

Vision-language models carry their own failure signatures. Engineering analysis has documented repetition loops and recitation errors as structurally distinct problems, each with its own root cause, its own API behavior, and its own mitigation path, and both have traced back to production outages. Open-source parsers show a recognizable pattern too: misordered columns in complex tables, dropped footnotes that anchor critical fields, and lost context across page breaks. OmniDocBench flags footnotes and figure or table captions specifically as sources of reading-order ambiguity extend.ai.

What ties all of this together is that none of it triggers an obvious exception. A mortgage system misreads a loan amount. An insurance system drops a coverage exclusion entirely. Every one of those is a silent failure, not a crash, and that's precisely why field-level measurement catches errors that uptime monitoring here would let pass unnoticed. RAG systems and fine-tuned LLMs that receive malformed or incomplete document parses produce unreliable outputs regardless of model quality. The parsing layer is the bottleneck most teams discover last, usually after something has already gone wrong downstream.

Building a per-field evaluation framework on your own documents

Start by defining the field schema before measuring anything. List every field the system needs to extract, assign each one a type (exact-match, numeric, date, free-text, table cell), and specify what class of consequence follows if that particular field comes back wrong. A wrong policy number and a wrong marketing description are not the same kind of error, and the evaluation framework should treat them differently from the start.

Building the ground-truth set is the next piece, and it has to be built on real documents, not demo PDFs, covering the actual range of layouts, scan quality, and edge cases the pipeline will meet in production. Use two independent annotators per document, and measure how often they agree with each other, since that agreement rate establishes the practical ceiling for that field type; the 83 to 91 mean AP range documented for footnote boundaries is a useful reference point for what "good agreement" looks like Pharma Document Extraction Benchmark: Tables & Footnotes | IntuitionLabs. Size the sample deliberately: 200 records for a ±4 percent margin at 90 percent confidence is the documented approach from the Donohue et al. study, and it's a reasonable starting point, though fields where errors are costlier deserve a larger sample Structured dataset of reported cloud seeding activities in the United States (2000-2025) using an LLM extend.ai.

Metric choice per field type isn't a stylistic preference, it changes what the number actually means. Exact-match rate suits identifiers, amounts, dates, and codes, where there is exactly one correct answer. CER handles OCR-sensitive text fields where character-level errors matter. TEDS is built for table cells and structured grids. Semantic similarity, whether through normalized information distance or vector-based comparison, fits free-text fields where paraphrasing is acceptable and exact wording isn't the point.

Confidence thresholding is where a lot of otherwise careful setups quietly go wrong. Avoid calibrating the threshold on the same split used to fit the confidence score itself, since that's exactly the score-refit leakage failure mode described earlier, and it inflates apparent performance in a way that doesn't survive contact with new documents. A fit/val split is the minimum standard, with a held-out set that never touches threshold selection at any point.

Measurement isn't a one-time project, either. Academic benchmarks test against fixed datasets, but production document distributions drift constantly, as vendors change invoice layouts, forms get revised, and new document types enter the pipeline. Automated regression testing against a growing evaluation set catches degradation before it reaches production at scale, and feeding human corrections back into that evaluation set does double duty: it fixes the immediate error and it improves the measurement baseline for everything that comes after. What should get reported, in the end, is per-field accuracy broken out by field type, precision and recall reported separately, the error rate among accepted fields at whatever confidence threshold is in use, and the trend over time, never a single aggregate figure standing in for all of it. Accounting for document clustering when sizing calibration sets, the design effect of 1.84–2.45 (per arXiv:2608.14639) means you need more documents than naive field counts suggest digitalapplied.com Valid Per-Field Selective Risk Control for Document Extraction: Three Failure Modes, a Validity Ladder, and When Conditioning Pays.

The build-vs-buy decision as a measurement problem

Building an extraction system in-house means building the measurement infrastructure too, and that second part is easy to underestimate until it's already late. Ground-truth sets, calibration pipelines, confidence scoring, and regression harnesses all have to be constructed alongside the extraction system itself, not after it ships.

The open-weight landscape in 2026 is genuinely strong on raw capability. None of these ship with a managed API, a service-level agreement, or a validation UI attached digitalapplied.com extend.ai Valid Per-Field Selective Risk Control for Document Extraction: Three Failure Modes, a Validity Ladder, and When Conditioning Pays.

Self-hosting starts to pencil out economically somewhere around 50,000 to 100,000 pages a month, and only with a dedicated ML engineer on staff to run it koreadeep.com digitalapplied.com. Below that volume, the engineering overhead erodes whatever per-page savings looked attractive on a spreadsheet koreadeep.com digitalapplied.com. Compute, storage, and egress for a self-hosted parsing setup typically run $15,000 to $40,000 a year even at modest scale, and maintaining parser dependencies, handling version conflicts, and writing custom logic for edge cases eats up 20 to 30 percent of an engineer's working hours per quarter extend.ai digitalapplied.com. Building means every field-level accuracy regression becomes the team's own problem to detect and fix, with no contractual floor on the error rate being absorbed in the meantime.

Commercial pricing gives a useful point of comparison for what "buy" costs. Azure AI Document Intelligence runs $1.50 per 1,000 pages for basic OCR, $10 per 1,000 for prebuilt models, and $30 per 1,000 for custom extraction, with a 500-page monthly free tier koreadeep.com checkthat.ai. Accuracy varies across providers even on a shared benchmark: on RD-TableBench, a set of 1,000 complex tables, one provider reports 90.2 percent average table accuracy against Azure Document Intelligence's 82.7 percent. None of these numbers settle whether to build or buy by themselves. What they do is turn it into a question that gets answered field by field, against ground truth, not by a single number on a vendor's homepage. Mistral OCR 4 costs $4 per 1,000 pages, with bounding boxes, block types, and confidence scores, undercutting Azure's custom tier by up to 15x digitalapplied.com.

Sources

  1. Valid Per-Field Selective Risk Control for Document Extraction:Three Failure Modes, a Validity Ladder, and When Conditioning Pays
  2. Structured dataset of reported cloud seeding activities in the United States (2000-2025) using an LLM
  3. Pharma Document Extraction Benchmark: Tables & Footnotes | IntuitionLabs

More in Accuracy Measurement