The obvious difference is the technology. AI on the saw reads images and controls machines, and AI in the office reads text and records. The more useful difference is what each is checked against. A machine grader has an obvious known answer, the grade a trained inspector would give, and published validations test it against one. Office AI has known answers too, the figures finance has already signed off, but nothing in the way it is sold asks for that test. Our recommendation is to give office AI the kind of test machine grading gets.
What the usual answer says
Comparisons of the two describe the technology, or its effect on jobs. A 2025 central bank working paper by Gustavo de Souza, "Artificial Intelligence in the Office and the Factory", found that "AI displaces routine office tasks while making machines more productive and easier to operate". That is a finding about employment. It says nothing about how far either kind of AI should be trusted.
How a machine grader is tested
A 2018 validation of an automated hardwood grader, published by Rado Gazo and colleagues, shows the shape of the test. More than 1,000 boards of each of nine species were graded by the machine and checked by a trained inspector. The result was reported two ways: "Across the entire volume of boards scanned, the automated grading system was 99.50% on-value and 92.22% on-grade accurate". Title 22 covers that study in detail. What matters here is the design: a known answer, a large sample and a reported rate.
It also shows that saw errors are not always visible. About one board in thirteen got a different grade from the inspector's, and only the inspector could see it. A mis-cut board is obvious. A misgraded one needs a check against a known answer.
Why office AI needs the same test
Office errors are harder still to see. A wrong margin figure looks like a right one, sits in a report next to correct figures, and travels into decisions before anyone checks it. It is found when someone reconciles it against another record, if anyone does.
The error rates are not small for at least one kind of office AI. In a paired benchmark published in April 2026 by researchers at Cube, which sells a semantic layer, three models answered analytical questions over one retail database. From the schema alone they got between 45.5% and 50.5% right. With a file of business definitions added, they reached 67.7% to 68.7%, so about a third of answers still failed. Document readers, forecasts and drafting tools will have different rates, which is the reason to measure your own.
An acceptance test for the office
Take a reporting tool. The known answers are figures finance has already signed off, and the sample is a set of questions with those answers. Set a pass mark before the tool is used, and re-run the test when the data or the model changes. Keep the questions that fail, because each one names a rule or a definition the business has never written down. For a document reader, the known answers are documents checked by hand, and the sample should include the worst scans as well as the clean ones.
One difference from the saw is real. A board has one correct grade under a written rule. An office question such as margin by customer can have several defensible answers, depending on how freight or rebates are treated. So the office test measures agreement with the business's own convention, and writing that convention down is part of the work.
The two also meet. A saw optimiser chooses between products using values that someone in the office keeps current, as title 43 sets out. Some of the most useful office AI may be the kind that checks those values against what the grades actually sold for.
When it doesn't apply
Office tools that only draft text, such as a first version of an email that a person always rewrites, need no test beyond the person reading it. Where no known answer exists, such as a forecast of next quarter, the test has to be run afterwards. Compare the forecasts with what happened.
Quarri for sawmills is built around how a sawmill runs, from log intake to shipped order.
Sources
- de Souza, "Artificial Intelligence in the Office and the Factory: Evidence from Administrative Software Registry Data", Federal Reserve Bank of Chicago Working Paper 2025-11: ideas.repec.org
- Gazo, Wells, Krs and Benes, "Validation of automated hardwood lumber grading system", Computers and Electronics in Agriculture, December 2018 (abstract read via Europe PMC): doi.org
- Rumiantsau and Fokeev, "Semantic Layers for Reliable LLM-Powered Data Analytics", arXiv 2604.25149, 28 April 2026: arxiv.org
Quarri is an AI-native data platform for the timber supply chain. It connects buying, production, sales and inventory for forest management, sawmill, wood products and pulp, paper and packaging operators.