Data lineage is the record of where data came from and what was done to it on the way: which sources fed it, which steps transformed it, and when they ran. It can be kept at three levels. At the table level, it shows which systems and jobs fed a report. At the field level, it shows which columns and which definition produced a figure. At the record level, it shows which individual records, and which documents, sit behind a number. Standard lineage models work mainly at the first two. For timber data, where numbers come from tickets, settlements and invoices, the third level is usually the one that matters.
What the usual answer says
Standard definitions describe lineage as tracking data's origins, movements and transformations through its life, shown in lineage maps, for governance, compliance, debugging and trust. Wikipedia's entry also separates coarse-grain lineage, of jobs and files, from fine-grain lineage, down to the data point. That is accurate. The illustrations are usually boxes for systems and arrows for pipelines, which is the table level.
The table and field levels
Lineage tooling has standardised around the pipeline view. OpenLineage, an open standard, describes itself as "an open framework for data lineage collection and analysis." Its core model defines "a generic model of dataset, job, and run entities": the data, the process that transformed it, and each time that process ran. That answers questions like which job produced this table, and when did it last run?
The field level adds meaning. Jiaqi Yin and colleagues, in an August 2025 preprint on enterprise pipelines, describe how complex transformations across several programming languages "often cause a semantic disconnect between original metadata and downstream data." They call it "semantic drift", which "compromises data reproducibility and governance". Their benchmark holds 1,700 hand-annotated lineages from industrial scripts, capturing source schemas, source tables, transformation logic and aggregation. They tested 12 language models on extracting it, and found performance rose with model size and more careful prompting. Their concern is metadata lost across pipeline scripts. The timber version, which is our framing rather than their finding, is a field called volume in one system that means something different in another.
The record level
A manager looking at stumpage paid on a tract needs more than the pipeline. They need the settlement lines that were summed, and the pages those lines were read from. That is record-level lineage, sometimes called provenance. It is heavier to keep, because it links every figure to many records. It is also the level at which a person can check a number against paper most directly.
For timber businesses, much of the data starts as documents: scale tickets, trip tickets, settlements, contracts, invoices. Record-level lineage has to reach back to those documents, beyond the table they were loaded into.
Lineage and AI answers
AI changes the question. A person asking an assistant how much did we pay this contractor last quarter? gets a number, with no pipeline in sight. Yiming Lin, Sepanta Zeighami and Aditya Parameswaran, in an August 2026 paper, put the problem plainly: language models return "answers to queries on data, without any indication for where the answer came from or whether it is trustworthy." Heuristics such as asking the model or matching similar text "could provide some hints for where the answer was derived", but "they provide no guarantees that the answer can be derived using the identified provenance, and indeed, are often incorrect". Their method, in a preprint tested across seven datasets, works after the fact: it finds a small set of source records and checks that they reproduce the answer, with over 30% higher accuracy than the best baseline.
The lesson is that a guessed source is not lineage. Provenance for an AI answer should either be recorded when the answer is computed or verified by re-deriving the answer from it. In our view, recording it at compute time is simpler where the system chooses the records first and then answers from them.
How much lineage to keep
Match the level to the use. Table-level lineage is enough to debug a failed pipeline. Field-level lineage is needed wherever definitions matter, such as margin, volume and cost. Record and document-level lineage is needed for any number people act on or report: payments, accruals, inventory values, investor figures. Our piece on proving where a number came from sets out what that record needs to hold.
Lineage or reconciliation?
Record-level lineage for every number is expensive, and much of the trust in a finance team's figures comes from reconciliation instead: does the report tie to the ledger, and the ledger to the settlement statements? If totals reconcile, clicking through to individual tickets is an audit feature, used when something is questioned. The two work together. Reconciliation says a total is right; record-level lineage says which records make it up when someone needs to check one.
When it doesn't apply
One-off analyses no one will need to reproduce can skip lineage beyond a note of the source. Small businesses running reports straight from one system get table-level lineage for free, because there is only one source. And record-level lineage adds storage and complexity, so it belongs where the numbers carry money or decisions.
How Quarri works explains the platform as a layer over existing systems, not a migration.
Sources
- OpenLineage documentation: openlineage.io
- Yin, Chen, Lee, Liu et al., "Schema Lineage Extraction at Scale: Multilingual Pipelines, Composite Evaluation, and Language-Model Benchmarks", arXiv 2508.07179, 10 August 2025: arxiv.org
- Lin, Zeighami and Parameswaran, "Bolt-on, Verifiable Provenance for LLM-Powered Data Processing", arXiv 2608.25210, 25 August 2026: arxiv.org
Quarri is an AI-native data platform for the timber supply chain. It connects buying, production, sales and inventory for forest management, sawmill, wood products and pulp, paper and packaging operators.