Yes, and some of the costly ones are confident. A July 2026 study of small open models on financial questions found that when a model gave the same answer on all eight separate attempts, that answer was still wrong 15 to 23% of the time. Asking the model again, or asking it to check its own work, does little against an error it makes consistently. For a business using a hosted tool, what catches it are checks outside the model: a control total the figure must agree with, a recalculation done in code or a spreadsheet, and a count of the records the figure was built from.
What the usual answer says
The advice that ranks for this question is to double-check AI output, ask the model to verify its answer, and keep a person in the loop. One consultancy's guide goes further and says to "Route all arithmetic to a Python function, spreadsheet API, or dedicated math service". That is sound, and it is one of the checks below. Asking the model to verify leans on the model itself, and a person in the loop needs to know what to check.
Why asking again doesn't work
Richard Zhe Wang published "Confidently Wrong" in July 2026, testing how well confident errors can be detected in financial question answering. The paper starts from the risk: "Hedged, uncertain answers invite scrutiny, whereas confident errors silently degrade downstream decisions without warning." It measured confidence as agreement across eight resampled answers to the same question. The finding: "among confident answers, those for which all eight resamples agree, 15-23% are wrong on FinQA". The models tested were small open models of eight or nine billion parameters, without tools, and part of that rate may be errors in the benchmark's own answers. No source here measures the same thing for frontier models with code execution. The paper also found that probes reading the model's internal activations detect these errors better than the model's own self-assessment, which it recommends as "a cost-effective triage mechanism". A business using a hosted model can't read its activations, which leaves checks on the output.
A separate 2026 paper by Hao Chen and colleagues, on numerical hallucinations in financial question answering, names three problems that retrieval-based systems face: "noise sensitivity, calculation fragility, and an auditability crisis". Its remedy is to have the model write a program that computes the answer, rather than state the answer, so the calculation can be run and checked. That points to the checks a business can apply itself.
Three checks outside the model
The first is a control total. Most timber figures have a total they must agree with. Delivered volume by contract should add up to the scale's total for the period. Margin by customer should add up to the gross margin in the accounts. If the AI's breakdown does not sum to the known total, something is missing or double-counted.
The second is a recalculation. Have the figure computed from the records by a query or a spreadsheet formula, rather than stated by the model in prose. Arithmetic done in code is repeatable and can be inspected, though units and rounding still need checking.
The third is a record count. Know how many records a figure should use: loads in the month, invoices for the customer, tags in the yard. Compare it with the number the AI's answer was built from. A correct calculation over an incomplete set of records gives a confident, consistent, wrong answer.
Where to put the checks
Checks done by hand after the fact get skipped when time is short. For each figure a business relies on, write down once its control total and its expected record count, so the checks can run wherever the figure is produced, and a figure that fails is marked rather than passed on.
The checks catch human errors too
These checks were not invented for AI. From Quarri's own work with a sawmill: a join error was hiding about half of finished inventory from reporting. No AI made that error. A record count against the physical inventory, or a control total against the ledger, is the kind of check that exposes an error like it, whoever or whatever produced the figure.
When it doesn't apply
Rough estimates, used only to decide whether a question is worth pursuing, don't need all three checks. Figures read directly from a single record, such as the total on one invoice, need only a glance at the record. And where no control total exists, as with a forecast, the check has to wait until the actual figure arrives. The forecast should be compared with it then.
Quarri for finance and strategy teams is built for the people who close the month, explain the margin and answer the board.
Sources
- Wang, "Confidently Wrong: Detecting Hallucinations in Financial Question Answering from LLM Internal States", arXiv 2607.11414, 13 July 2026: arxiv.org
- Chen et al., "Fighting Numerical Hallucinations via Data-centric Compilation for Online Financial QA", arXiv 2605.31064, 29 May 2026: arxiv.org
- Dojo Labs, "Why AI Gets Math Wrong and How to Actually Fix It": dojolabs.co
- Quarri evidence ledger, E15 (proven)
Quarri is an AI-native data platform for the timber supply chain. It connects buying, production, sales and inventory for forest management, sawmill, wood products and pulp, paper and packaging operators.