AI in forestry and lumber · 28 Sep 2026

What data do you need before AI is useful?

← AI in forestry and lumber

It depends on the question, and usually less than a data programme assumes. For any one question, AI needs the records the answer is built from, enough history to cover the question's own cycle, and keys that join those records, such as a contract number on both the load and the invoice. The cycle might be a day, a quarter or several seasons. A business can have all of that for one question and none of it for the next. Asked about data as a whole, leaders give contradictory answers, and a survey published in January 2026 shows how.

What the usual answer says

The advice that ranks for this question says AI needs large amounts of clean, complete, consistent data, and that it should be put right first. One guide says the right data "depends on your specific needs". None of them says how much history a question needs, or which keys have to join. That leaves a business with no way to know when it is ready.

Question Records it needs History it needs Key that joins Daily tally One day's records One day Date and line Margin by customer Every invoice line, carrying customer and product in the same form The period the business wants to compare Customer and product 15-day harvest forecast Daily harvest records Several seasons, each covered more than once Not stated
Readiness belongs to a question, not to the data as a whole. A daily tally is ready as soon as its records are digital; a seasonal forecast needs seasons that no cleaning can bring forward. Diagram: Quarri.

The readiness paradox

A survey by Drexel University's LeBow College and Precisely, fielded in 2025 and published in January 2026, asked 505 data and analytics leaders at large companies about readiness. It found that leaders "confidently report having the necessary infrastructure (87%), skills (86%), and data readiness (88%) for AI, but a large proportion also admit infrastructure (42%), skills (41%), and data readiness (43%) are their biggest obstacles". By simple arithmetic, at least three in ten respondents must have given both answers on data. The report also warns that "This data quality debt raises substantial risk if not addressed". Precisely sells data integrity software and has an interest in the finding, and the respondents work at larger companies than most timber businesses.

In our reading, the contradiction makes sense once readiness is tied to a question. A business's data can be ready for the questions it already answers in spreadsheets and not for the ones it has never tried.

What one question needs

Take a harvest forecast. From Quarri's own work with a forestry operation: a 15-day harvest forecast, trained on four years of daily data, was cross-validated to within about 260 tonnes a day. Our reasoning is that several years matter for a question like that because they cover each season more than once.

Research points the same way from the other side. A 2024 study in Forests by Almeida, da Silva and Simões built two datasets for predicting harvester productivity, "both with 25,387 instances": one of harvest records alone, one with weather added. Three of the weather features were removed as having "low relevance". The features that mattered most were two already in the harvest records: "working hours and average individual tree volume". Adding a data source added little. The records the question needed did the work.

Questions and their requirements

A daily tally needs one day's records, joined by date and line, and is ready as soon as those records are digital. Margin by customer needs every invoice line to carry the customer and the product in the same form, over whatever period the business wants to compare. A seasonal forecast needs several seasons recorded, and no amount of cleaning brings that date forward.

How to check readiness

Write the question down. List the records its answer is built from and where each lives. Say how far back the answer needs to look. Name the key that joins each pair of records, and test it on a sample: how many records carry it, and how many match.

One caution runs the other way. If each question gets its own keys and definitions, a business ends up with several answers to "what is our volume". The keys and definitions a first question uncovers should be agreed once and reused, so the second and third questions start from them. That is the useful part of a data programme, and it can grow one question at a time.

When it doesn't apply

Exploratory work, which looks for questions worth asking, cannot list its records in advance, and a broader cleanup may be the only way to start. Some questions need data a business has never collected, and there the first step is recording it.

How Quarri works explains the platform as a layer over existing systems, not a migration.

Sources

  1. Drexel LeBow Center for Applied AI and Business Analytics and Precisely, "2026 State of Data Integrity and AI Readiness", survey fielded in 2025, published January 2026: lebow.drexel.edu
  2. Almeida, da Silva and Simões, "Cut-to-Length Harvesting Prediction Tool: Machine Learning Model Based on Harvest and Weather Features", Forests 15(8), 1398, 10 August 2024: doi.org (read via mdpi-res.com)
  3. RWS, "Determining what AI data you need and how to source it": rws.com
  4. Quarri evidence ledger, E19 (proven)

Quarri is an AI-native data platform for the timber supply chain. It connects buying, production, sales and inventory for forest management, sawmill, wood products and pulp, paper and packaging operators.

See it on your own data.

Live in two weeks, on the systems you already run.