It should say so. The field, or the whole document, should come back marked unread, with a reason: illegible, page missing, format not recognised, no matching contract, arithmetic doesn't add up. The document goes to a queue for a person, and the queue is watched for age. What should never happen is an unchecked guess: a plausible number in the field where the smudge was, posted with nothing to test it. And the reasons, counted over months, show which senders and forms cause most of the trouble, which is where the fix belongs.
What the usual answer says
The usual answer describes an exception workflow. Documents the AI can't handle go to a human for review and manual entry, and the system learns from the corrections. That describes what happens after a tool admits it can't read something. It assumes the tool will admit it.
Why tools guess instead
Research published in September 2025 by Adam Tauman Kalai and colleagues, three of them at OpenAI, a model developer, Why Language Models Hallucinate, sets out the cause. "Language models hallucinate because the training and evaluation procedures reward guessing over acknowledging uncertainty." Most tests score a blank and a wrong answer the same, so a model that guesses scores better. In the authors' words, "guessing when uncertain improves test performance".
Document reading shows the same pattern. Zhentao He and colleagues, in June 2025, found that models reading degraded documents such as invoices often invent content "especially when a precise answer is not feasible". When they trained a model to recognise uncertainty and refuse in those cases, its hallucination-free accuracy rose 22 points above a leading general model on their own benchmark. Refusing is a skill, and it has to be asked for.
Exceptions are normal
Every document process has them, and unreadable pages are only one kind. Ardent Partners surveyed 212 accounts payable professionals, most from large companies, for its 2025 metrics report, sponsored by the e-invoicing company Pagero. Its exceptions are "Invoices flagged due to coding errors, missing information, approval bottlenecks, lack of purchase order data" and similar, and not documents a machine couldn't read. It put the exception rate at 14% in 2024, with 9.0% for the best performers against 22.0% for the rest. And for the first time in the study's history, "invoice exceptions (53%) sit as the top challenge in the AP industry". An unreadable ticket joins a queue that is already busy.
So the aim is not to eliminate exceptions. It is to make sure every one is visible, and none is hidden inside a guess.
How to design for it
Make unread an outcome. Every field can be returned as read, unread or doubtful, with a reason code. A document with an unread field does not post.
Score tools the right way. When testing an extraction tool, count a wrong value as worse than a blank; we use three times worse as a starting point. A tool that guesses less and admits more will score better.
Set the line by what can be checked. Abstention has a cost: every blank is a person's time, and a tool that refuses too readily floods the queue. Where a field has a check behind it, such as arithmetic or a match to the scale record, a doubtful value can go through and let the check decide. Where nothing can test it, a blank is safer.
Age the queue. Track how long each unread document has waited, and set a limit by document type, shorter for settlements and invoices.
Use the reasons. Count exceptions by reason and by sender each month. In our experience a few senders or forms cause a large share: a contractor who faxes, a buyer whose statements print too small, a trip ticket with a field in the fold. Ask those senders for a better copy or a digital file, or change the form. A fix at the source removes that cause, where model training only works around it.
Watch one number above the rest: exceptions cleared by a person who entered the same value the tool had marked doubtful. A high share means the tool is too cautious and the line can move. A low share means its doubts are real.
What an exception record should hold
Keep it short and specific. The document and page. The field that failed. The reason code. The sender. The date it arrived and the date it was cleared. Who cleared it, and what value they entered.
That last pair matters most. The value a person enters becomes a checked example. Over time, those examples show whether a tool is improving on the senders that trouble it. They also give an honest accuracy figure for the hard cases, which vendor figures rarely cover.
When it doesn't apply
Documents that arrive as structured data or born-digital PDFs rarely need reading at all, so exceptions are about missing data rather than legibility. Very low volumes can go straight to a person. And some documents are unreadable by anyone, such as a torn or water-damaged ticket, where the right outcome is a request for a copy, not better extraction.
Quarri for operations teams is built for the people running the crews, the lines and the yard.
Sources
- Kalai, Nachum, Vempala and Zhang, "Why Language Models Hallucinate", arXiv 2509.04664, 4 September 2025: arxiv.org
- He, Zhang, Wu, Chen et al., "Seeing is Believing? Mitigating OCR Hallucinations in Multimodal Large Language Models", arXiv 2506.20168, 25 June 2025: arxiv.org
- Ardent Partners, "Accounts Payable Metrics That Matter in 2025": datocms-assets.com
Quarri is an AI-native data platform for the timber supply chain. It connects buying, production, sales and inventory for forest management, sawmill, wood products and pulp, paper and packaging operators.