Documents and paperwork · 28 Sep 2026

How do you check that AI read a document correctly?

← Documents and paperwork

Mostly with checks that don't depend on a person spotting the error. First, the document's own arithmetic: net against gross less tare, lines times rates against the total. Second, a match to another record: the ticket against the load in the scale system, the settlement against the contract. Third, where a person does review, show each extracted value next to the part of the image it came from, so they check the source rather than the answer. Fourth, audit the checks themselves now and then, by planting known errors and counting how many are caught.

What the usual answer says

The usual advice is to use confidence scores, send low-confidence fields to a person, add validation rules and spot-check against the original. All sensible. It treats human review as the safety net. The evidence says the net has holes, and that how the review is set up decides how big they are.

Arithmetic Lines times rates that don't sum to the total; net that isn't gross less tare Record match Tonnage that doesn't match the scale tickets the settlement covers Grounded review A wrong figure that happens to fit, checked beside the crop of the page Planted-error audit Reviewers approving rather than reading; the real catch rate
Each layer passes on less for the next. The first two need nobody to notice anything; review starts from the page, not the answer; and planted errors measure whether the whole process is catching what it should. Diagram: Quarri.

Why human review is weaker than it looks

Jacob Beck and colleagues ran a randomised experiment with 2,784 participants, published in September 2025, in which people reviewed AI suggestions on an annotation task. Two findings matter for anyone checking extracted documents. First, "requiring corrections for flagged AI errors reduced engagement and increased the tendency to accept incorrect suggestions." Making review more work made it worse. Second, "individual attitudes toward AI emerged as the strongest predictor of performance, surpassing demographic factors". "Participants skeptical of AI detected errors more reliably and achieved higher accuracy, while those favorable toward automation exhibited dangerous overreliance on algorithmic suggestions."

The authors conclude that success depends on the algorithm and equally on "who reviews AI outputs and how review processes are structured". The participants were a crowd doing an annotation task, and trained clerks reviewing their own suppliers' documents may do better. Our inference is that the direction holds: a reviewer looking at a clean extracted figure is inclined to accept it.

Why confidence thresholds are not enough

Confidence scores decide which fields a person sees, so they deserve their own check. ConfBench, published in August 2026 by Priyashree Roy and colleagues at Amazon Web Services, which sells document extraction, tested confidence estimates on 1,346 degraded document variants and more than 70,000 field-level evaluations. It found calibration "varies widely across models, from near-perfect to severely overconfident". With an overconfident model, the wrong fields never reach review at all. The paper also finds a remedy: "per-model post-hoc correction rescales these absolute confidence values for threshold-based routing". So a threshold can be made to work, but it has to be set for the model in use. We would recalibrate on a sample of your own documents whenever the tool's model changes.

Checks that don't need anyone to notice

Arithmetic. Many timber documents carry figures that must agree. Net weight equals gross less tare. Lines times rates sum to the total. Deductions match the contract terms. A field that breaks the arithmetic is flagged whatever the confidence score says.

Record matches. The document should agree with something else the business holds. A ticket matches a load at the scale. A settlement's tonnage matches the tickets it covers. An invoice matches a purchase order and a receipt. A document with no match goes to a queue, not the ledger.

These two catch many wrong numbers without relying on anyone's attention, though not a wrong figure that happens to fit, which is what the next two layers are for.

Review that starts from the source

When a person does review, change what they look at. Instead of a table of extracted values, show each value beside the crop of the page it was read from. Research on extraction is moving this way. STNet, published in 2024 by Shuhang Liu and colleagues, is designed to "deliver precise answers with relevant vision grounding", pointing to the region of the image behind each answer. Whatever tool is used, the principle is to make the source the first thing the reviewer sees.

Keep review light. Beck's finding suggests that piling corrections onto reviewers reduces care. Send them only what the mechanical checks couldn't settle.

Audit the checks

We suggest a regular audit, quarterly for most volumes. Take a sample of documents that passed every check, including review, and plant known errors in a copy: a changed weight, a swapped product code, a missing line. Run them through the same process and count what gets caught. That rate is the real measure of the checking, and it will show whether reviewers are reading or approving.

When it doesn't apply

Documents with no internal arithmetic and nothing to match against, such as the special conditions in a contract, rely on careful reading, and the reviewer should read the original, not the summary. Very low volumes may be checked in full by a person, which makes planted errors less necessary. And where extraction feeds only a search index, not the ledger, a lighter process is reasonable.

Quarri for operations teams is built for the people running the crews, the lines and the yard.

Sources

  1. Beck, Eckman, Kern and Kreuter, "Bias in the Loop: How Humans Evaluate AI-Generated Suggestions", arXiv 2509.08514, 10 September 2025: arxiv.org
  2. Roy, Martin, Rostami, Romo et al., "Can You Trust the Confidence? ConfBench for Vision-Language Models on Document Extraction", arXiv 2608.01792, 3 August 2026: arxiv.org
  3. Liu, Zhang, Hu, Ma et al., "See then Tell: Enhancing Key Information Extraction with Vision Grounding", arXiv 2409.19573, 29 September 2024: arxiv.org

Quarri is an AI-native data platform for the timber supply chain. It connects buying, production, sales and inventory for forest management, sawmill, wood products and pulp, paper and packaging operators.

See it on your own data.

Live in two weeks, on the systems you already run.