Research through a data lens · 28 Sep 2026

How accurately does document AI read handwritten forms?

← Research through a data lens

Close to a careful person on free text, less well on handwritten tables, and wrong in ways that are harder to spot. An August 2026 benchmark tested 18 models on 500 real handwritten documents. The best scored 71.85 on a composite of text, table and formula measures, against 77.09 for people transcribing the same pages in one pass. It beat them on text and trailed them by over 14 points on tables. More of its errors were plausible substitutions: 66.56% of its table errors, against 34.48% of the humans'.

What the usual answer says

The usual answer gives 80 to 99% accuracy for handwritten forms, depending on legibility, with human review of low-confidence fields. Vendor pages at the top of the search repeat figures in that range, usually without saying what was counted or how. The research is more useful because it says both.

Overall score composite of text, table and formula Best model 71.85 Humans, one pass 77.09 Text: the best model ahead Tables: humans ahead by over 14 points Table errors that were prior-driven fluent output the image did not support Best model 66.56% Humans, one pass 34.48%
The best model came within a modest gap of people overall, but its table errors were more often fluent substitutions the page did not support, which are harder to spot. Source: WildHandBench, August 2026. Diagram: Quarri.

What the 2026 benchmark found

WildHandBench, by Zhang and colleagues (arXiv 2608.22959), covers free text, tables and formulas from nine real-world settings, forms among them. Its documents are 66.6% Simplified Chinese and 28.6% English. The authors say it "should not be interpreted as a population-level estimate of handwritten OCR performance".

The 71.85 is an average of three scores: a text edit-distance score, a table similarity score and a formula score. It is not a share of fields read correctly. The paper contrasts it with 96.34% for the top model on a separate printed-document benchmark, a different set of pages.

Against people, "the relatively modest gap indicates that wild handwriting challenges humans and models alike". On text, "the best model actually surpasses humans". Tables were the weak point. The best table score "indicates nearly 40% structural and content mismatch", and humans led on tables by more than 14 points.

The authors also classified each error by whether it came from the image or from the model's expectations. Models tend toward "generating fluent but visually unsupported output where humans tend toward conservative partial transcriptions". They add a caution: a high prior-driven rate "does not mean a model is less reliable, only that its errors are more systematic". Compact, OCR-focused models had the lowest rates and made more visible misreadings.

A test on real forms

Businessware Technologies, a document-processing developer with an interest in the result, published a January 2026 benchmark scored field by field on 10 real hand-filled forms. A field counted as correct at 80% text similarity, but was marked wrong when errors hit names, dates, IDs or phone numbers. Its conclusion: "Even the best models rarely exceed 95% business-level accuracy". Ten forms is a small sample, and the company doesn't publish per-field results.

The 2025 study

A 2025 benchmark by Crosilla, Klic and Colavizza found general models "achieve excellent results in recognizing modern handwriting", mostly clean lines of European-language text. It also found that the models, the 2024 generation, "demonstrate limited ability to autonomously correct errors in zero-shot transcriptions". Clean modern lines are easier than degraded pages with ruled tables, which is how the two benchmarks can both be right.

Reading it through a data lens

Timber forms are mostly tables of numbers, which is the harder case in the 2026 results. The error type matters. A garbled field gets noticed. A 7 read as a 1 in a column of diameters, or a volume that fits the rows above, does not. For numbers, a plain similarity threshold is too lenient: one wrong digit in a five-digit number still scores 80%.

In our view, the useful checks sit outside the page. Columns should add to the written totals and stay within physical ranges. Better still, each form should agree with another record of the same load, such as its weighbridge weight.

The strongest objection is that timber forms are the easy case: labelled fields, digits only, a small vocabulary and known ranges. That may be right. Neither benchmark tested timber forms, so the only way to know is to key a sample of a business's own sheets by hand and compare field by field, scoring numbers exactly.

When it doesn't apply

Forms typed or filled in on screen sit in the printed range, where accuracy is far higher. In our judgement, tick boxes and single-digit boxes are easier than free handwriting or ruled tables.

Quarri for operations teams is built for the people running the crews, the lines and the yard.

Sources

  1. Zhang, Zhao, Cui, Qu, Sun, Yang, Zhou, Liu and Han, "WildHandBench: A Benchmark for Handwritten Text Understanding that Challenges MLLMs and Humans", arXiv 2608.22959, 24 August 2026, read in full: arxiv.org
  2. Businessware Technologies, "Handwritten Form Recognition Benchmark: Accuracy, Cost, and Performance Comparison of Leading AI Models", January 2026: businesswaretech.com
  3. Crosilla, Klic and Colavizza, "Benchmarking Large Language Models for Handwritten Text Recognition", arXiv 2503.15195, revised 23 June 2025 (abstract): arxiv.org

Quarri is an AI-native data platform for the timber supply chain. It connects buying, production, sales and inventory for forest management, sawmill, wood products and pulp, paper and packaging operators.

See it on your own data.

Live in two weeks, on the systems you already run.