Data practice · 28 Sep 2026

How clean does data need to be before AI is useful?

← Data practice

Clean enough for the question being asked, and for most business questions that means free of systematic error in the fields the question uses. Scattered errors, such as a misspelt customer or an impossible board length, rarely move a total. Systematic errors, repeated across many rows in the same direction, move every total built on them: one supplier's lines in the wrong unit, records that never close, a missing month, a join that drops records. Those are also the errors automated cleaning, including AI cleaning, handles worst. For payments and invoices, where each record is acted on, scattered errors matter too.

What the usual answer says

The top answers now agree that cleanliness is relative. "There's no such thing as 'clean data,'" one data leader told CIO in November 2024. "It's always relative to what it is you're using it for." That is right, and it stops businesses waiting for perfect data. It doesn't say which errors to look for.

Row-level errors Systematic errors Scattered; visible in one row Repeated across many rows, one direction A misspelt customer name An impossible board length Units wrong on one supplier's lines Records that never close A missing month A join that drops records Rarely moves a total Leave to checks as they come up Moves every total built on it And hardest for automated cleaning to catch Between the two: one supplier, product or tract under several codes
For most business questions, the errors worth cleaning first are the systematic ones. They move totals, and they only show when rows are compared. Diagram: Quarri.

What the research shows

The research is on cleaning data for training machine learning models, a different setting from AI answering business questions; the transfer is our judgement. Tommaso Bendinelli and colleagues, in a March 2025 workshop study on deliberately corrupted public datasets, found language models "can identify and correct erroneous entries, such as illogical values or outlier", using other fields in the same row. But "they struggle to detect more complex errors that require understanding data distribution across multiple rows, such as trends and biases".

REIN, a 2023 benchmark by Mohamed Abdelaal and colleagues at a data software company, tested 38 error detection and repair methods on 14 datasets to ask "where and whether data cleaning is a necessary step in ML pipelines". It found some cleaning strategies "cause the predictive performance to be substantially deteriorated", so cleaning isn't automatically an improvement. A 2023 review by Pierre-Olivier Côté and colleagues catalogues the field across 101 papers from 2016 to 2022, including entity matching as a cleaning task of its own.

A systematic error, in practice

From Quarri's own work with a lumber and millwork manufacturer: a units labelling error had inflated recorded purchase spend by about a fifth.

No single line looked impossible. The error only shows when lines are compared, which is the cross-row kind the research says AI cleaning struggles with.

Broken identity

A third kind of error sits between the two: the same supplier, product or tract under different codes in different systems. Each record is valid, and no distribution looks odd, but an assistant asked for spend with one supplier will give half the answer. For question-answering, matching those identities is often the first cleaning job.

How to decide what to clean

Start from the question and list the fields it depends on. Then test those fields for the systematic errors that usually matter in timber data. Records that never close, such as tags, open orders or accruals, show up in an age profile. A unit or conversion that is wrong for one supplier shows up in implied prices per unit. Missing periods show up as gaps in a monthly count, and joins that drop or duplicate records show up when totals are compared before and after the join. Fix those, and match entities that appear under several codes. For totals, leave scattered row errors to checks as they come up.

Where AI does help

The research finding is about AI asked to clean a table unprompted. Asked a specific question, the same tools can compute the profiles that expose cross-row errors: the age profile of open records, implied price per unit by supplier and month, record counts by period, totals before and after each join. The person still has to know which profiles to ask for. Once the checks exist, they can run every month, so a systematic error is caught in the period it starts rather than at year end.

Why waiting costs more

Cleaning without a question has no finish line, because there is always another field. It also tends to fix what is easy to see, the row errors, while the systematic ones stay hidden until a question exposes them.

When it doesn't apply

Training a model from scratch, such as a grading or defect model, needs clean labels across the training set. Payments, invoices and regulatory reports need every record right, as well as the total. And some businesses have data so poor that no question can be answered, in which case capture, not cleaning, is the first job.

Quarri for operations teams is built for the people running the crews, the lines and the yard.

Sources

  1. Bendinelli, Dox and Holz, "Exploring LLM Agents for Cleaning Tabular Machine Learning Datasets", arXiv 2503.06664, 9 March 2025: arxiv.org
  2. Abdelaal, Hammacher and Schoening, "REIN: A Comprehensive Benchmark Framework for Data Cleaning Methods in ML Pipelines", arXiv 2302.04702, 9 February 2023: arxiv.org
  3. Côté, Nikanjam, Ahmed, Humeniuk et al., "Data Cleaning and Machine Learning: A Systematic Literature Review", arXiv 2310.01765, 3 October 2023: arxiv.org
  4. CIO, "When is data too clean to be useful for enterprise AI?", 27 November 2024: cio.com
  5. Quarri evidence ledger, E6 (proven)

Quarri is an AI-native data platform for the timber supply chain. It connects buying, production, sales and inventory for forest management, sawmill, wood products and pulp, paper and packaging operators.

See it on your own data.

Live in two weeks, on the systems you already run.