AI in forestry and lumber · 28 Sep 2026

Can you ask your ERP questions in plain English?

← AI in forestry and lumber

Yes. Current AI models have become much better at turning a plain-English question into a database query. Whether the answer is right depends on something most ERPs don't hold: a written record of what their fields mean. In an April 2026 benchmark, three frontier models got about half of 99 analytical questions wrong from the database schema alone. A four-kilobyte document of business definitions raised their accuracy by 17 to 23 percentage points.

What the usual answer says

The vendor pages that rank for this question describe the experience. An ERP vendor's explainer promises that you can "ask your business data a plain question and get a straight answer". Another product promises that "Warehouse and operations staff run your ERP by talking to their phone or pointing the camera at it." The better pages add a warning about data quality: "If one customer is saved under three slightly different names, your sales answer will be wrong, and it will sound confident while being wrong."

Share of 99 questions answered correctly Claude Opus 4.7, Claude Sonnet 4.6 and GPT-5.4, same database Schema only 45.5% to 50.5% Schema plus a definitions file of about 4 KB 67.7% to 68.7% The gap between models was too small to separate statistically in either condition.
Adding a short written record of what the data means lifted all three models by more than switching model could. The outlined end of each bar is the spread across the three models. Source: a 2026 benchmark by Cube, which sells a semantic layer. Diagram: Quarri.

The same explainer adds that "A good answer uses your definitions", so the commodity view does name meaning. What none of these pages gives is a measurement: how often answers are wrong, and how much of that comes from the query and how much from what the data means.

How good the models have become at writing queries

The standard test for enterprise-scale query writing is Spider 2.0, presented at ICLR 2025. Its databases come from real applications, "often containing over 1,000 columns". When the paper appeared in late 2024, its own agent, built on OpenAI's o1-preview, solved 21.3% of the full tasks, against 91.2% on the older academic benchmark, Spider 1.0. The simpler Spider 2.0-Lite set, with prepared metadata and documentation, gives a like-for-like trend. On its public leaderboard, read on 25 September 2026, the same o1-preview agent scores 23.03 and the top system scores 76.23.

That is fast progress in under two years, and it is the part vendors are selling. It still leaves about a quarter of the Lite tasks failed on query writing alone. It measures whether a model can find the right tables and write correct SQL for a question whose meaning is fixed in advance. A question put to an ERP by a finance lead rarely arrives with its meaning fixed.

Where the wrong answers come from

Michael Rumiantsau and Ivan Fokeev isolated that second problem in a paired benchmark published in April 2026. They put the same questions to Claude Opus 4.7, Claude Sonnet 4.6 and GPT-5.4 twice over a retail sales database. The first time, each model saw only the schema. The second time, it also saw a hand-written file of about 4 KB. The file listed the measures and their formulas, the data's quirks and the rules for ambiguous terms.

Without the file, the three models scored between 45.5% and 50.5%. With it, they scored between 67.7% and 68.7%. The gap between models was too small to separate statistically in either condition. A team with the document can use whichever of the three models suits its budget, the authors write, and "A data team that has not cannot recover the accuracy gap by switching to a stronger or more expensive model."

Their account of why is the useful part for an ERP. "Many real analytical datasets have conventions a schema does not reveal," they write. One of their examples is a snapshot table whose values must not be summed across time. Their ambiguity example is "sales by region": customer region or store region? An ERP is full of fields like these. A quantity column may hold board feet on one line and pieces on the next. An order may carry a ship date and an invoice date. Last month's sales is then a different total depending on which one the query uses. The data can be clean and the answer still wrong, because nobody told the model which reading was meant.

It is one retail dataset of 99 scored questions. The answers were judged by a model from the same family as two of those tested. The authors are at Cube, which sells a semantic layer, and the benchmark is published under its GitHub account. The file was written by an analyst who had seen the dataset, and the authors note that "a truly blind document, written before any benchmark interaction, might encode less of the right knowledge". Even with the file, about a third of answers still failed.

The file is also the expensive part. On a clean retail warehouse it took 4 KB. An ERP with years of custom fields and site-specific habits will need more, written by the people who know what each field means. An ERP vendor's own semantic model, where one exists, may cover part of it.

What definitions can't fix

A second 2026 paper, by Morris Lee, lists failures that query-matching benchmarks never score. One is "queries that execute successfully while returning the wrong business number." Wrong numbers of that kind come partly from meaning and partly from the records themselves.

From Quarri's own work with a lumber and millwork manufacturer: a units labelling error had inflated recorded purchase spend by about a fifth. A definitions file tells the model which column holds purchase spend. It cannot tell it the column is wrong. Ask that ERP in plain English what the business spent with a supplier, and it will return the inflated figure, fluently and quickly. Catching that takes a check against a second record, such as the supplier's invoices, which a chat window over one system does not make.

A test before you trust the answers

Take 20 questions finance has already answered this quarter, with figures somebody has signed off. Ask each one in plain English and count how many tie exactly. For every miss, write down the reason: the wrong table, the wrong definition, the wrong period, or a figure in the ERP that is itself wrong.

The first three kinds of miss are fixed by writing the rule down, and that list is the definitions document the benchmark measured. The fourth kind is a data problem, and no model or document fixes it.

When it doesn't apply

Single-record lookups, such as an order's status or one customer's open balance, carry little ambiguity, and plain-English access to them works today with little preparation. The test matters for totals, margins and comparisons across periods. It also stops short where the answer sits partly outside the ERP, in a scale system or a production tally. A model reading one system answers a narrower question than the one asked.

Quarri for finance and strategy teams is built for the people who close the month, explain the margin and answer the board.

Sources

  1. Rumiantsau and Fokeev, "Semantic Layers for Reliable LLM-Powered Data Analytics: A Paired Benchmark of Accuracy and Hallucination Across Three Frontier Models", arXiv 2604.25149, 28 April 2026: arxiv.org (full text: arxiv.org)
  2. Lei et al., "Spider 2.0: Evaluating Language Models on Real-World Enterprise Text-to-SQL Workflows", ICLR 2025, arXiv 2411.07763 (v2, 17 March 2025): arxiv.org
  3. Spider 2.0 leaderboard, read 25 September 2026: spider2-sql.github.io
  4. Lee, "Business Truth, not SQL Accuracy: A Rule-Gated 7B Analytics Agent Outperforms a Direct-Prompted 32B Baseline", arXiv 2608.09254, 10 August 2026: arxiv.org
  5. Enterpryze, "How AI Insights in ERP Turn Your Data Into Answers": enterpryze.com
  6. ERPChat product page: erpchat.app
  7. Quarri evidence ledger, E6 (proven)

Quarri is an AI-native data platform for the timber supply chain. It connects buying, production, sales and inventory for forest management, sawmill, wood products and pulp, paper and packaging operators.

See it on your own data.

Live in two weeks, on the systems you already run.