AI in forestry and lumber · 28 Sep 2026

What is an AI agent, and what could one do in a lumber business?

← AI in forestry and lumber

An AI agent is a language model given tools and a goal, which it works towards in steps. It reads records, decides what to do next, uses a tool to act, and checks the result before the next step. For a lumber business, doing a task once matters less than doing it the same way the hundredth time. That points its first jobs at work with many repetitive records, where it reads and drafts and a person approves anything that changes a record or reaches a supplier.

What the usual answer says

The descriptions that rank for this question present agents as autonomous: software that orders stock, answers customers and runs schedules on its own. Vendors are building the connections to make that possible. Microsoft's documentation for its ERP analytics, for example, describes an MCP server that "enables AI agents to access and analyze Business performance analytics data through natural language". The server is in preview. Once agents have access like that, what matters is how often they get things right.

The agent reads and drafts Read open orders Every open purchase order and its promised date Compare Promised dates against receipts and history Draft A supplier message and an updated expected date Control point A person approves The message and the date Record updates Only after approval Next morning, the loop runs again Every step reads records the business already has. The only writes pass through a person.
The agent reads and drafts; a person approves anything that changes a record or reaches a supplier. That approval is the control point. Diagram: Quarri.

How reliable agents are

A useful measurement comes from τ-bench, published in June 2024 by Shunyu Yao and colleagues. It sets agents tasks in customer-service settings with purpose-built API tools, a simulated customer and written rules, and judges success by whether the database ends up in the right state. It also introduced a measure called pass^k: whether an agent succeeds on the same task in every one of k attempts.

In 2024 the authors found that "even state-of-the-art function calling agents (like gpt-4o) succeed on <50% of the tasks, and are quite inconsistent (pass^8 <25% in retail)". Models have improved. A June 2025 follow-up, τ²-bench, reported one newer model succeeding on 74% of retail tasks and 56% of airline tasks at the first attempt. It also found that scores "decline more rapidly" as the number of repeated attempts rises in its harder domain. We have not seen a published pass^k for the newest models on these tasks, so the gap may be narrower now. The measure is still the one to ask for: an agent that succeeds most of the time, but not on every run of the same task, is fine for drafting and risky for acting.

A business can run the same test on its own work before trusting an agent with it. Give the agent the same twenty records ten times over. Count how often its output agrees with itself across the runs, and how often it agrees with the answer a person would give. Both counts matter, and a single demonstration shows neither.

A task that suits an agent

Purchase orders suit an agent because every one carries a promised date and, later, a receipt, so each draft can be checked against a record. From Quarri's own work with a lumber and millwork manufacturer: under 4 in 10 purchase orders arrived on time, across 13,000+ closed orders, and late ones averaged about nine days.

An agent can read every open order each morning. It compares promised dates with receipts and the supplier's history, and flags the orders at risk. Then it drafts two things: a message to the supplier, and an updated expected date for the planner. A person approves the message and the date. Each step reads records the business already has. The only writes, the message and the date change, pass through a person.

Once the agent's drafts have been checked over many runs and found consistent, the approval can move from every draft to a sample.

Where to be careful

Tasks that change prices, release orders or commit money should stay with a person longer. A wrong action there is expensive to undo, and a reliability short of every time is not enough. Tasks that read and summarise, such as a morning list of at-risk orders or a check of invoices against receipts, carry little risk from an occasional miss.

When it doesn't apply

One-off questions, such as a single margin analysis, need no agent. A person working with an AI assistant is simpler and easier to check. A business with few purchase orders or customers has little repetitive work to hand over. And an agent is only as good as the records it reads. If promised dates are not recorded, no agent can chase them.

Quarri for wood products is built around how a wood products plant runs, where the order book meets real capacity.

Sources

  1. Yao, Shinn, Razavi and Narasimhan, "τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains", arXiv 2406.12045, 17 June 2024: arxiv.org
  2. Barres, Dong, Ray, Si and Narasimhan, "τ²-Bench: Evaluating Conversational Agents in a Dual-Control Environment", arXiv 2506.07982, 9 June 2025: arxiv.org
  3. Microsoft Learn, "What is Business performance analytics?", updated July 2026: learn.microsoft.com
  4. Quarri evidence ledger, E1 (proven)

Quarri is an AI-native data platform for the timber supply chain. It connects buying, production, sales and inventory for forest management, sawmill, wood products and pulp, paper and packaging operators.

See it on your own data.

Live in two weeks, on the systems you already run.