Documents and paperwork · 28 Sep 2026

How accurate is AI at reading invoices and settlement statements?

← Documents and paperwork

Per field, often very accurate on clean documents from familiar senders. Per document, less so, because a document is only right if every field is, and a settlement statement has many. Every load, rate, deduction and adjustment is a field. Recent benchmarks, from researchers at document AI vendors, also show performance falling as schemas get wider and pages get worse. The honest answer for any business is a figure measured on its own documents, per document and per line, with arithmetic checks to catch what gets through.

What the usual answer says

Vendors and guides put AI invoice extraction at 95% to 99% accuracy on common fields, well above manual keying, and say it improves with use. The better ones add a caveat. Parseur, which sells extraction, says "Roughly 99% is real on clean invoices from repeat suppliers" and "It is not real on every field of every document your suppliers will ever send". They are per-field figures, usually on the easiest fields. They are not the chance that a whole document is right.

Chance a whole document is right, if field errors are independent 0% All about 67% about 30% about 74% 20 60 Short invoices Settlement statement, 60 line fields Fields per document 99.5% per field 98% per field 95% per field
Per-field accuracy compounds. A long settlement statement is where it bites; errors that cluster make real results better than this, which only a per-document measure shows. Diagram: Quarri.

Why per-field accuracy misleads

Errors compound. If each field is right 98% of the time, independently, a document with 20 fields is entirely right about 67% of the time, since 0.98 multiplied by itself 20 times is about 0.67. A settlement statement with 60 line fields at the same rate is entirely right about 30% of the time. At 99.5% per field, the 60-field statement comes out right about 74% of the time. Independence is the pessimistic case. In practice errors cluster: a bad page produces several wrong fields while clean pages produce none, so more documents come out entirely right than the arithmetic predicts. How much better depends on how the errors cluster in your documents, which only a per-document measure shows. The direction still holds: the more fields, the less a per-field figure says.

That arithmetic is ours. The benchmarks point the same way, with limits.

What the benchmarks show

ExtractBench, published in February 2026 by Nick Ferguson and colleagues at Contextual AI, which sells document AI products, pairs 35 documents with detailed schemas and human-checked answers, giving 12,867 fields to score. Its finding on frontier models was blunt. They "remain unreliable on realistic schemas". Performance "degrades sharply with schema breadth, culminating in 0% valid output on a 369-field financial reporting schema across all tested models". Valid output there means parseable JSON that fits the schema, scored separately from correctness, and the documents are filings, credit agreements, research papers, résumés and sports results, not invoices. The paper also notes that "Valid JSON also does not imply correct extraction": on one domain, 90% of outputs were valid and 12.5% passed.

The benchmark also sets out why scoring is hard. Fields "demand different notions of correctness (exact match for identifiers, tolerance for quantities, semantic equivalence for names)", and "omission must be distinguished from hallucination". A missing deduction and an invented one are different errors with different fixes.

Confidence scores are the usual defence, sending uncertain fields to a person. ConfBench, published in August 2026 by researchers at Amazon Web Services, which sells document extraction, found calibration "varies widely across models, from near-perfect to severely overconfident". Its degraded variants are grouped by how far they cut accuracy, with the worst tier costing more than ten points. A confident wrong field goes straight through.

What this means for timber documents

Invoices from contractors and suppliers are short and mostly header fields, so per-field and per-document accuracy are close. Settlement statements, scale summaries and stumpage reports are long tables: many loads, each with weight, product, rate and amount. They are the documents where compounding bites, and where a single wrong line can move a payment.

Timber documents also carry their own checks. Lines times rates should sum to the total. Tonnage should match the tickets it covers. Deductions should match the contract. A document that fails its own arithmetic is flagged whatever the confidence scores say.

Where the errors tend to fall

ConfBench's degradation tiers support the first part of this: worse pages, lower accuracy. The rest is our view. We would expect errors to gather on the pages hardest to read, such as faxes, photographs and carbon copies, and on the fields that vary most between senders, such as how each buyer labels deductions. A tool that reads most buyers' settlements well can still stumble on one layout. So report accuracy by sender and document type, and not as one average.

How to measure it

Take a hundred recent documents of each main type. Key every field by hand, or use documents already reconciled. Run the extraction and score three things: the share of fields right, the share of lines entirely right, and the share of documents entirely right. The third figure is the one that says how many documents a person will need to touch. Then see how many of the wrong documents the arithmetic checks caught. The ones they missed are the real error rate.

When it doesn't apply

Short, clean, digital invoices with a handful of fields behave close to the per-field figures, and whole-document accuracy is high. Documents where the business only needs the total, not the lines, can be scored on that field alone. And a tool fine-tuned on one buyer's settlement format may do far better on it than general benchmarks suggest; measuring on your own documents shows it.

Quarri for finance and strategy teams is built for the people who close the month, explain the margin and answer the board.

Sources

  1. Ferguson, Pennington, Beghian, Mohan et al., "ExtractBench: A Benchmark and Evaluation Methodology for Complex Structured Extraction", arXiv 2602.12247, 12 February 2026: arxiv.org
  2. Parseur, "AI Invoice Processing Benchmarks 2026", last updated 8 September 2026: parseur.com
  3. Roy, Martin, Rostami, Romo et al., "Can You Trust the Confidence? ConfBench for Vision-Language Models on Document Extraction", arXiv 2608.01792, 3 August 2026: arxiv.org

Quarri is an AI-native data platform for the timber supply chain. It connects buying, production, sales and inventory for forest management, sawmill, wood products and pulp, paper and packaging operators.

See it on your own data.

Live in two weeks, on the systems you already run.