AI in forestry and lumber · 28 Sep 2026

Can AI be trusted with timber numbers?

← AI in forestry and lumber

Not as a system, only answer by answer. No published benchmark shows an AI tool accurate enough on financial figures to be trusted blind, and the best recent results still leave a large share of answers wrong. A figure can be trusted when it arrives with the records behind it and a check that every relevant record was counted. For a timber business that means the loads, tickets, rates and invoices behind each number, and a count showing none were left out. This piece is about business figures, such as volumes, costs and money owed. The measurement tools that rank for this question, which count and measure logs from photos, are a different test.

What the usual answer says

Vendors answer this question with conditions: clean data, a good tool and a human reviewing the output. One ERP vendor puts the principle well: "A good answer shows its working. If it says three customers are over their limit, you can see which three. You are never asked to just take its word".

Delivered volume for one contract, one month The answer Loads it counted Rate applied to each Load ticket, date, contract the rate the contract sets Load ticket, date, contract the rate the contract sets Load ticket, date, contract the rate the contract sets Load ticket, date, contract the rate the contract sets and every other load in the answer Loads counted against loads on the scale for that contract and month: no more, no fewer
A figure can be trusted when it arrives with the records behind it and a check that none were left out. The completeness line is the part that runs by itself. Diagram: Quarri.

That is the right principle. It needs a measure of how often answers without their working go wrong, and a definition of what the working should include.

How often answers go wrong

The clearest early measurement is FinanceBench, published in November 2023 by Pranab Islam and colleagues. It holds 10,231 questions about publicly traded companies, each with an answer and "evidence strings", the passages of the filing that support it. The authors tested 16 model configurations on a 150-question sample and reviewed 2,400 answers by hand. Their headline for one configuration: "GPT-4-Turbo used with a retrieval system incorrectly answered or refused to answer 81% of questions". They concluded that "all models examined exhibit weaknesses, such as hallucinations, that limit their suitability for use by enterprises."

Newer tests measure different things, so they aren't one trend line, but none has removed the problem. A 2025 benchmark of 537 expert-written financial research questions found that "even the best-performing model (OpenAI o3) achieved only 46.8% accuracy". A paired benchmark published in April 2026, by researchers at the semantic-layer vendor Cube, asked three frontier models 100 analytical questions against a database. A written file of business definitions raised accuracy by 17 to 23 points, and the models then answered between 67.7% and 68.7% correctly. What the model was given mattered more than which model it was.

What the working should include

FinanceBench's own design shows the remedy. Every answer in it has evidence attached, so a reviewer can see whether it is right without redoing the work. A timber figure needs the same: the records it counted, and the rule it applied to each.

It also needs something a filing question does not: a completeness check. A delivered-volume figure can be built from correct records and still be wrong because some loads were missed. From Quarri's own work with a forestry operation: one release cycle matched 1,400+ bills of lading across 70+ invoices, with every rated row and every eligible weigh record accounted for.

How to apply it

When an AI tool gives you a number that matters, ask what records it is built from, what rule was applied to each, and how many records there should have been against how many were used. A tool that can answer those for every figure can be trusted figure by figure. A tool that gives a number and a confident sentence cannot, however good its average accuracy.

Nobody will audit forty record lists a week, so the checks have to run by themselves. One way is to keep the model away from the arithmetic: it picks a defined, reconciled metric, and a fixed calculation produces the figure (the subject of title 33). The other is for the business to know, for its main figures, what count of records to expect, so a mismatch is flagged rather than found. That takes work once. For a monthly delivered volume by contract, the expected count is the number of loads on the scale for that contract and month. The check is that the AI's figure used exactly those loads, no more and no fewer, and that each carried the rate the contract sets.

When it doesn't apply

Exploratory questions, where an approximate answer is enough to decide whether to look further, don't need full evidence each time. Some figures come straight from a single record, such as one invoice's total, and there the check is trivial. And a human-built spreadsheet deserves the same test. It is not made trustworthy by being built by hand, only by showing its records.

Quarri for finance and strategy teams is built for the people who close the month, explain the margin and answer the board.

Sources

  1. Islam et al., "FinanceBench: A New Benchmark for Financial Question Answering", arXiv 2311.11944, 20 November 2023: arxiv.org
  2. Rumiantsau and Fokeev, "Semantic Layers for Reliable LLM-Powered Data Analytics", arXiv 2604.25149, 28 April 2026: arxiv.org
  3. Bigeard and others, "Finance Agent Benchmark: Benchmarking LLMs on Real-world Financial Research Tasks", arXiv 2508.00828, 20 May 2025: arxiv.org
  4. Enterpryze, "How AI Insights in ERP Turn Your Data Into Answers", undated: enterpryze.com
  5. Quarri evidence ledger, E13 (proven)

Quarri is an AI-native data platform for the timber supply chain. It connects buying, production, sales and inventory for forest management, sawmill, wood products and pulp, paper and packaging operators.

See it on your own data.

Live in two weeks, on the systems you already run.