Research through a data lens · 28 Sep 2026

Can harvester data predict productivity? What the studies found

← Research through a data lens

Yes, to a point. A 2022 study trained 24 machine learning algorithms on harvester data to predict productivity. Its best model reached a coefficient of determination of 0.71, with a mean absolute percentage error of 15%. A 2024 study from the same group, adding weather, reached 0.70. A 2023 study found average tree size had the greatest effect on productivity. The 2022 abstract adds a detail it doesn't explain. The dataset held 144,973 records collected over 28 months, and "After cleaning the initial database, we used only 1.12% to build the model." We read all three from their abstracts; the publisher's full texts weren't reachable.

Source: Munis, Almeida, Camargo, da Silva, Wojciechowski and Simões, "Machine Learning Methods to Estimate Productivity of Harvesters", Forests 13(7): 1068, published 7 July 2022, open access, doi.org/10.3390/f13071068.

What the usual answer says

The top results include a 2020 review of research on harvester on-board computer data and the 2022 study itself. The usual answer is that harvester data predicts productivity from tree size, terrain, operator and machine data with good accuracy for planning. The studies support the accuracy. They say less about the data behind it.

2022 study: the records behind the model 144,973 records over 28 months After cleaning 1.12% used to build the model roughly 1,600 records reason not stated in the abstract The model R² of 0.71 coefficient of determination MAPE of 15% mean absolute percentage error 2023 study Average tree size had the greatest effect
The 2022 model's accuracy describes the 1.12% of records left after cleaning, and the abstract does not say why the rest were set aside. The 2023 study found tree size had the greatest effect on productivity. Diagram: Quarri.

What the studies did

The 2022 study aimed "to analyze the performance of machine learning algorithms in estimating the productivity of timber harvesters". The predictors were "the availability of hours of machine use, individual mean volumes of trees, and terrain slopes". The team tested "24 machine learning algorithms in default mode", then blending and stacking. "Learning by blending ensemble stood out with a determination coefficient of 0.71 and a mean absolute percentage error of 15%". The 1.12% is roughly 1,600 records.

The 2024 study combined "15 weather and timber harvesting attributes", tested 24 algorithms again, and a tuned model reached "a coefficient of determination of 0.70". Its abstract doesn't say whether the data matches the 2022 set, so the two figures can't be read as showing that weather added nothing.

A 2023 study used data from a manufacturer's cloud service for two mid-sized harvesters, with "as little processing as possible". Linear regression showed "average tree size was the variable having the greatest effect on fuel consumption per cubic meter and productivity". Productivity averaged "around 10 m3/h". It notes that the standard harvester files most researchers use "are troublesome to obtain and require some pre-processing". The cloud data was easier to get, but using it as it came limited the authors to unsupervised methods. They write that "Extending the database in follow-up studies will facilitate the application of supervised learning techniques for modeling and prediction."

Reading it through a data lens

The 1.12% is the most useful number in the 2022 abstract, and the least explained. It could mean aggregation, with many time records rolled into shift or stand observations. It could mean filtering to one machine type or regime, or removing bad records. Each would change what the model describes: all the work, a subset, or the clean part of it. The abstract doesn't say which. A model's accuracy should be reported with the share of records used and the reason.

The 2023 finding points the same way. If tree size explains most of productivity, then clean, stand-linked tree volumes matter more than more variables. Anyone offered a productivity model built on harvester data can ask three things. What share of the records did it use? Why were the others dropped or merged? And does it predict the stands and crews that were left out?

A worked check

Say a contractor exports a year of harvester records and finds 60,000 rows. It checks each for a stand ID, an operator, a shift and a known stop category. If 9,000 pass, it knows its usable share before building anything, and it knows why the rest failed. The figures are illustrative.

What we cannot take from it

We read abstracts, so the cleaning criteria, the unit of a record and the study areas' detail are unknown to us. The prediction studies come from plantation systems that differ from mixed natural forests. An R² of 0.7 still leaves much of the variation unexplained, so predictions can guide plans rather than individual days.

Quarri for forest management is built around how a forest operation runs, from the cruise to the settled account.

Sources

  1. Munis, Almeida, Camargo, da Silva, Wojciechowski and Simões, "Machine Learning Methods to Estimate Productivity of Harvesters", Forests 13(7): 1068, 2022 (abstract via Crossref): doi.org
  2. Almeida, da Silva and Simões, "Cut-to-Length Harvesting Prediction Tool: Machine Learning Model Based on Harvest and Weather Features", Forests 15(8): 1398, 2024 (abstract via Crossref): doi.org
  3. Polowy and Molińska-Glura, "Data Mining in the Analysis of Tree Harvester Performance Based on Automatically Collected Data", Forests 14(1): 165, 2023 (abstract via Crossref): doi.org

Quarri is an AI-native data platform for the timber supply chain. It connects buying, production, sales and inventory for forest management, sawmill, wood products and pulp, paper and packaging operators.

See it on your own data.

Live in two weeks, on the systems you already run.