Yes, for most tickets, and the misreads to worry about are the ones that look right. A benchmark of handwritten documents published in August 2026 found that most model errors come from the model's expectations rather than the page. The model writes down what a document of that kind usually says. On a scale slip, that error is a plausible weight. Confidence flags catch the smudged field, so the checks that matter are the ticket's own arithmetic and a second record of the same load.
What the usual answer says
The vendor pages at the top of the results agree that it works. A guide from Lido says most teams see eight or nine in ten fields extract cleanly. It adds that "Lido uses confidence scoring to flag uncertain extractions for human review rather than silently guessing." A forestry ticketing vendor argues for replacing paper altogether, and promises, without published evidence, to "Reduce admin time by 50%+".
Both answers rest on a model knowing when it is unsure. The recent evidence says that for handwriting it often does not.
What the benchmark found
WildHandBench, published in August 2026 by Zhang and colleagues, tested 18 current models on 500 handwritten documents, including handwritten tables, against calibrated human readers. The best model scored 71.85% overall and humans 77.09%. On printed documents, the authors note, the top model on a leading benchmark "now reaches 96.34% overall". Handwriting is still a different problem.
The finding that matters for tickets is about the kind of error. The authors built a measure of whether an error came from the image or from the model's language priors. Between 63 and 91% of model errors were prior-driven, against 49% for humans. In their words, "humans remain conservative on illegible content, whereas models confidently hallucinate fluent but unsupported text."
The paper also finds the gap is narrow, and that top models beat humans on text transcription under some metrics. The difference is in how each fails. Carried over to a scale house, which is our extrapolation, a clerk who cannot read a weight is more likely to leave a query, and a model more likely to supply a number that fits. On a ticket full of figures, a fitted number is indistinguishable from a read one unless something else disagrees with it. The benchmark has clear limits for this use. Two thirds of its 500 documents are in one language other than English, most are prose, and only 81 are tables, so it does not measure tickets directly. A digit in a weight field gives a model far less to guess from than a sentence, so prior-driven errors may be rarer on tickets. What the benchmark shows is the direction of the error, and that direction is the one a confidence score is least suited to catch.
Checks a ticket already supports
A trip ticket and a scale slip carry their own redundancy. Gross less tare should equal net. A load count should match the lines. A date should fall within the haul period of the contract it cites. We expect a misread of one figure to break one of these relations more often than not. But a model reading the whole ticket at once can also produce a gross, tare and net that agree with each other and are all wrong, and the arithmetic check misses that. It is cheap to run on every ticket once the fields are structured, and it catches some errors a reviewer scanning flagged fields would pass.
That is why the stronger check is a second record. The same load usually exists somewhere else: the scale's own weigh record, the mill's receiving entry, the settlement line that pays for it. A read ticket that ties to one of those is confirmed. One that does not tie goes to a person, whatever the model's confidence. The same tie catches a failure the benchmark does not measure at all: a ticket read correctly and attached to the wrong load.
From Quarri's own work with a forestry operation: a pilot turned 20 handwritten trip tickets into structured load records. That is pilot scale, and it shows the reading is feasible. It does not show an error rate across thousands of tickets, and no figure here should be read that way.
When it doesn't apply
Printed scale slips from a scale-house system are a different case. The printed-document figure the authors cite, from a different benchmark, suggests reading them is close to solved, and the risk shifts to matching them to the right load. Where a ticket is the only record of a load, the arithmetic check is all there is, and a person should review a larger share. The benchmark also says nothing about a specific model on a specific operation's handwriting, so run a sample of your own tickets against the checks before relying on it.
Quarri for forest management is built around how a forest operation runs, from the cruise to the settled account.
Sources
- Zhang et al., "WildHandBench: A Benchmark for Handwritten Text Understanding that Challenges MLLMs and Humans", arXiv 2608.22959, 24 August 2026: arxiv.org
- Lido, "How to Extract Data from Handwritten Documents with AI", 8 July 2026: lido.app
- Waldo, "The Real Implications of Switching to Digital Trip Tickets in Forestry", 10 December 2025: waldologs.com
- MemX, "Can AI Read Your Handwriting? Mostly" (read as a top result, not cited): memx.app
- Quarri evidence ledger, E23 (pilot)
Quarri is an AI-native data platform for the timber supply chain. It connects buying, production, sales and inventory for forest management, sawmill, wood products and pulp, paper and packaging operators.