Join each internal record to the public figure as it stood when the record was made, on the same unit, definition and price level. For most aggregated price indexes, revisions are small and the first release will do. For detailed indexes, rail data and thin price reports, store every release with its date and join by that date. Then compare changes as well as levels. The step most often missed is timing. Public series change after release, and AI models may already know what happened next. Either can make a comparison look sharper than it was.
What the usual answer says
The usual answer is about tools and alignment: load both sets into one database or dashboard, standardise units and periods, match categories, and chart the variances. That is necessary. The time problem gets less attention.
Public figures change after release
The statistics agency's notes on the producer price index say that since November 2021, "indexes undergo five iterative updates from first issuance through final posting", with final indexes four months after first publication. The same page is reassuring about size. First releases "typically are based on a substantial portion of the total number of responses", so "subsequent revisions normally are minor, especially at more highly aggregated grouping levels". Seasonally adjusted indexes are recalculated each year, going back five years.
Other series move more. The railroads' weekly traffic summary says its data "are subject to revision for up to a year".
So the question is when vintage discipline is worth the effort. For a contract escalator, the answer is simple: use the release the contract names, and the revision is irrelevant. For a review of buying decisions against a broad index, the first release and the final will usually tell the same story. The cases that matter are detailed or thin series, series revised for many months, and seasonally adjusted figures compared across years. In those, a purchase made in March and judged against a figure published in July is being judged against a number no one had in March.
Point-in-time joins
The fix is well established in data tools. Keep each public figure with the date it was released, and join each internal record to the latest version released on or before its own date. The pandas documentation describes this as a merge "similar to a left-join except that we match on nearest key rather than equal keys". Its backward search "selects the last row in the right DataFrame whose 'on' key is less than or equal to the left's key". Most databases support the same pattern.
You may not need to store the vintages yourself. The ALFRED archive, run by a regional central bank, "allows you to retrieve each economic data release (vintage) that was available on a specific date in history". Its home page lists more than 25,000 price series, and its page for the lumber producer price index lists releases from March 2015 to September 2026. Rail carloads, stumpage reports and trade association surveys generally have no such archive, which is where keeping your own dated copies pays off.
With vintages available, a business can ask two questions cleanly. How did our buying compare with the market as we saw it then? And how does it compare with the market as we now know it was?
The AI version of the same problem
AI brings a new form of the same trap. Language models are trained on text up to a cutoff date, and that text includes what happened to prices and markets. Asked to explain or forecast a past period, a model may use knowledge it could not have had at the time.
A December 2025 study by Gao, Jiang and Yan built a test for this and applied it to forecasts of stock returns from news headlines and of capital expenditure from earnings calls. It measures the chance a model "has internalized information about the realized outcome". That chance was "materially positive throughout the in-sample period" and "collapses essentially to zero right after the training-data cutoff". Forecasts looked more accurate where the model already knew the answer. Timber prices weren't tested, but the mechanism is the same.
Researchers are responding with models trained only on text available up to each date. A 2026 paper by Kelly, Malamud, Schwab and Xu describes "point-in-time language models" with up to 4 billion parameters, trained on "1 trillion chronologically filtered tokens" with monthly checkpoints from 2013 to 2024. They narrow the gap with open models of similar size trained without date limits, though a gap remains on several tasks. The authors note that models trained on unrestricted text "inevitably embed information from the future". The practical test for any AI analysis of past periods is to check how it performs on periods after its training data ends.
When it doesn't apply
Purely descriptive comparisons of settled history, where no decision is being judged and the public series has stopped changing, can use final figures. So can most comparisons against broad, aggregated indexes, given how small their revisions normally are.
Quarri for finance and strategy teams is built for the people who close the month, explain the margin and answer the board.
Sources
- Bureau of Labor Statistics, Handbook of Methods, "Producer Price Indexes: Presentation": bls.gov
- Association of American Railroads, "Weekly Railroad Traffic" summary, week 37, 2026: aar.org
- pandas documentation, "pandas.merge_asof": pandas.pydata.org
- Federal Reserve Bank of St. Louis, ALFRED: alfred.stlouisfed.org and lumber PPI series page: alfred.stlouisfed.org
- Gao, Jiang and Yan, "Detecting Lookahead Bias in LLM Forecasts", arXiv, 29 December 2025: arxiv.org
- Kelly, Malamud, Schwab and Xu, "Scaling Point-in-Time Language Models", arXiv, 2026: arxiv.org
Quarri is an AI-native data platform for the timber supply chain. It connects buying, production, sales and inventory for forest management, sawmill, wood products and pulp, paper and packaging operators.