Because a scan is a picture of the data. It can be viewed, sent and stored, and a searchable PDF can be searched for a word. But the numbers on it can't be added up, joined to other records or checked against them until they are read out into fields, such as tract, species, volume and price, with units and keys. The part most businesses haven't done is the archive. Documents scanned years ago have usually never been set beside the ledger, and that comparison is where reading scans into data pays.
What the usual answer says
The top answers to this question already make the distinction. A scan is an image; OCR makes it searchable; digitisation turns it into data that, in one guide's words, can "sum amounts, check for mistakes". That is right. It treats digitisation as something done to documents as they arrive, and says nothing about the documents already scanned and filed.
A searchable scan is not a check
A searchable PDF carries a hidden OCR text layer, and its quality is unknown. A 2024 benchmark, OHRBench, tested OCR solutions, including traditional pipelines and vision-language models, on 8,561 document images with tables, formulas and charts. It found "none is competent for constructing high-quality knowledge bases" for AI retrieval systems. It tested knowledge bases for AI rather than keyword search. Still, it shows that text read from images needs checking before anyone relies on the numbers in it. Fields can be checked: a settlement's lines times rates should equal its total, and its volume should match the scale tickets it covers.
What reading the archive finds
From Quarri's own work with a timberland manager: 140+ timber sale documents were read with every scanned page recovered, and the owner's own ledger turned out to be under-counting volume on some properties.
The difference showed once the documents' figures became fields beside the ledger's. We think this is the general case for archives, though one proof point doesn't show how common it is. A scanned document is checked when someone looks for it; fields can be checked against every other record at once.
A difference is a finding, not yet a correction. The ledger may be wrong, or a document may be: a missing amendment, a sale recorded twice, a volume later revised. Each difference needs the page it came from beside it, so someone can decide which record to change.
Archives add one problem current paperwork doesn't have: keys that changed. A tract renumbered after a boundary change, a buyer that merged, a contract numbering scheme replaced ten years ago. Fields read from old documents join to the ledger only through a mapping from old keys to new ones. Building that mapping is often most of the work, and the records that fail to map are worth a list of their own.
The cost of stopping at the scan
Documents that are only scanned still have to be found and read by a person whenever a number is needed. Atlassian, which sells collaboration software, surveyed 12,000 knowledge workers and 200 executives for its State of Teams 2025, and found that "leaders and teams waste 25% of their time just searching for answers". The figure covers searching in general. An archive of images is one of the places that searching happens.
How to go from scan to data
Decide first which documents carry numbers the business uses: sale contracts, settlements, scale tickets, invoices, cruise summaries. In our view, AI extraction has made it practical to include the archive as well as current paperwork, because the cost of reading a page is no longer a person's time at a keyboard. Extract the fields with their units, and keep the image beside every field as its evidence. Check each document against its own arithmetic, then against the ledger and the scale system. Expect differences on the first pass, and investigate each in both directions.
When it doesn't apply
Documents kept only as legal records, such as deeds and signed originals, need to be preserved and found, not computed, so a good scan is enough. Documents born digital, such as PDF invoices from accounting software, already carry exact text and need extraction but no reading of images.
Quarri for forest management is built around how a forest operation runs, from the cruise to the settled account.
Sources
- Zhang et al., "OCR Hinders RAG: Evaluating the Cascading Impact of OCR on Retrieval-Augmented Generation", arXiv 2412.02592, first posted 3 December 2024, v4 read (30 August 2025): arxiv.org
- Atlassian, "State of Teams 2025": atlassian.com
- Quarri evidence ledger, E22 (proven)
Quarri is an AI-native data platform for the timber supply chain. It connects buying, production, sales and inventory for forest management, sawmill, wood products and pulp, paper and packaging operators.