Software and systems · 28 Sep 2026

How do you evaluate a data platform for a timber business?

← Software and systems

On your own data, with questions whose answers you already know. Take last year's audited inventory value, a season's volume by tract, one supplier's annual spend, a month's overrun, and ask the platform for each. Score it exactly: right or wrong, not close. Add a few questions it should refuse, because the data isn't there. Then make one change, such as a new product code or a new levy rate, and see how long the platform takes to reflect it and who has to do the work.

What the usual answer says

Guides to choosing a data platform offer weighted scorecards: integrations, scalability, security, ease of use, cost, support and AI features, scored after a demo. Those criteria matter, and security and cost can rule a platform out. None of them measures whether it gives the right answer to a timber business's questions.

Section Example timber questions Pass rule Known-answer questions Inventory value at year end, volume by product and month, stumpage paid by tract Exactly right against the known figure, not close Questions to refuse A mill it has no records for, a year before the history begins Says it can't answer, instead of producing a figure Change test A new product code or a new levy rate Reflected correctly in days; note whose time it took Export test The data, the definitions, the rules and the keys Readable without the platform
The sheet runs on the business's own data. A platform that fails the first two sections should not reach the scorecard, whatever its features. Diagram: Quarri.

Why the demo isn't the test

BEAVER shows how far performance falls when data moves from public examples to real companies. It is a benchmark built by Peter Baile Chen and colleagues from private enterprise data warehouses, and holds 9,128 question and query pairs over 812 tables. Its authors note that language models do well on public benchmarks, while their efficacy in private enterprise environments, "characterized by intricate schemas, domain knowledge, and analytical user queries", "remains unproven". In their latest results, "SOTA agentic frameworks using the advanced model GPT-5.2 achieve only 10.8% accuracy." With expert hints for every sub-step, accuracy rose to 30.1%. Those are the features a vendor's sample data lacks and a timber business's data is full of.

BEAVER measures agentic systems answering cold against raw warehouses. It says nothing about a platform whose vendor has first modelled the business's data and definitions, so 10.8% is a floor for cold text-to-SQL, not a measured rate for any data platform. What it does show is that good answers on private data can't be assumed from good answers on public examples, and a demo on the vendor's data is a public example.

Known-answer questions

We suggest 25 to 30 questions whose answers you already hold, from sources you trust: audited statements, settled contracts, closed sales. Spread them across the business: stumpage paid by tract, volume by product and month, margin by customer for a closed year, inventory value at year end. Ask each in plain language, the way a manager would. Score each exactly against the known figure, and record whether the platform showed how it got there.

Known answers set a minimum bar. They don't prove the platform can answer what nobody has answered, which is often why it was bought. So add one or two questions of that kind that someone has worked out by hand once, such as a cross-system reconciliation of volume bought against volume consumed for one quarter.

Keep the set, and run it again after every upgrade and every change to the data. A platform that scored 27 out of 30 in the trial should still score 27 a year later. The set costs a few days to build and becomes the business's own benchmark, which no vendor can tune for.

Questions it should refuse

We suggest five questions the platform cannot answer from the data it holds: a mill it has no records for, a year before the history begins, a product the business doesn't sell. A good platform says it can't answer. A poor one produces a figure. A 2024 federal standards profile of generative AI risks calls this confabulation, and defines it as "the production of confidently stated but erroneous or false content". In our judgement it is the most damaging failure in an evaluation, because nothing marks it as wrong.

The change test

Pick one real change from the past year, apply it, and time how long the platform takes to reflect it correctly, and whose time it takes: yours, the vendor's, or a consultant's.

The export test

Finally, ask for everything back: the data, the definitions, the rules and the keys, in formats you can read without the platform. Note anything that can't be exported, since that is what a later move would have to rebuild.

Reading the results

A platform that scores well on known answers, refuses what it should, absorbs a change in days and exports cleanly is worth scoring on everything else. One that fails the first two should not reach the scorecard, whatever its features.

When it doesn't apply

Platforms bought purely for storage or backup, with no questions asked of them, need a different test. Very early evaluations, before any data is shared, can only compare features and terms. And known-answer tests need someone who knows the answers, which in a small business may be the same person running the evaluation.

How Quarri works explains the platform as a layer over existing systems, not a migration.

Sources

  1. Chen, Yang, Li, Wenz et al., "BEAVER: An Enterprise Benchmark for Text-to-SQL", arXiv 2409.02038, first posted 3 September 2024, latest version with 2026 results: arxiv.org
  2. National Institute of Standards and Technology, "Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile", NIST AI 600-1, July 2024: nvlpubs.nist.gov

Quarri is an AI-native data platform for the timber supply chain. It connects buying, production, sales and inventory for forest management, sawmill, wood products and pulp, paper and packaging operators.

See it on your own data.

Live in two weeks, on the systems you already run.