A deterministic system gives the same output every time it is given the same input. A spreadsheet formula is deterministic, and so is a saved database query. A language model, as usually served, is not, even at the setting meant to make it so. For financial data this matters because a figure that can change between two identical requests can't be reproduced for an auditor, a lender or a finance lead. The fix is to keep each reportable figure in a defined, stored calculation that the AI selects and explains. A query the AI writes fresh each time is only as repeatable as its text, so that text has to be saved and reused too.
What the usual answer says
The pages that rank for this question are mostly vendors arguing for governed, rule-based calculation. One analytics vendor's guide contrasts generative AI, whose output it describes as variable from run to run, with its own "Same answer every time" and "Full query traceability". That is the right direction. None of the top results shows how far a model's own output varies, or separates repeatable from correct.
What happens in practice
Horace He and colleagues at Thinking Machines Lab published a test in September 2025. They note that even at temperature zero, the setting meant to make a model pick its most likely word every time, "LLM APIs are still not deterministic in practice". They sampled 1,000 completions of a single prompt from a large open model at temperature zero. The result: "we generate 80 unique completions, with the most common of these occuring 78 times". The completions were identical for the first 102 tokens and then began to diverge.
The cause lies in how the inference software adds numbers. The order of additions changes with how requests are batched together on the server, and floating-point sums depend slightly on order. The authors show it can be fixed with batch-invariant settings, at a cost in speed, and open inference engines can offer that. Whether a hosted AI service exposes such a setting is a question to ask its provider.
A prompt about a historical figure is not a margin report, and a short numeric answer has fewer tokens in which to drift than a long essay. The mechanism is the same, though, and a figure written out in text could drift.
Repeatable is not the same as right
A deterministic model would still not make its numbers correct. A 2026 benchmark from dbt Labs, which sells a semantic layer, compared models writing their own SQL with models choosing from defined metrics. The authors wrote that with defined metrics the model "can't produce correct-looking numbers that are subtly different across runs : the logic is codified and deterministic". The same benchmark found accuracy higher with defined metrics, between 98.2% and 100% against 84.1% to 90.0% for SQL the models wrote themselves. The property a finance team needs is that each figure is defined and traceable. Repeatability follows from that.
The trade-off, in dbt's words, is coverage: defined metrics "can only answer questions that fall within the scope of what's been modeled". A question outside them needs a new definition, or an ad-hoc query that is saved with its answer.
What to ask for
A figure meant to tie to another record is where this shows. From Quarri's own work with a sawmill: a daily production tally that took about a quarter of an hour by hand now takes a couple of minutes, and ties to the dollar against the manager's own sheet. For any figure like that, the questions to ask of an AI tool are which defined calculation produced it, and whether running it again on the same records gives the same result.
When it doesn't apply
Drafting text, such as a summary of a contract or a first version of a supplier email, does not need determinism. Variation between drafts is harmless, and a person reads the result. Exploratory questions, where the aim is ideas rather than figures, don't need it either. And deterministic code can still be wrong, consistently, so it needs the same checks against control totals and record counts as any other figure.
How Quarri works explains the platform as a layer over existing systems, not a migration.
Sources
- He, Horace and Thinking Machines Lab, "Defeating Nondeterminism in LLM Inference", Connectionism, 10 September 2025: thinkingmachines.ai
- Ganz and Perigaud, "Semantic Layer vs. Text-to-SQL: 2026 Benchmark Update", dbt Labs, 7 April 2026: docs.getdbt.com
- Chata.ai, "The Complete Guide to Deterministic AI Analytics", 8 June 2026: chata.ai
- Quarri evidence ledger, E18 (proven)
Quarri is an AI-native data platform for the timber supply chain. It connects buying, production, sales and inventory for forest management, sawmill, wood products and pulp, paper and packaging operators.