Software and systems · 28 Sep 2026

What should an AI vendor tell you about your data before you sign?

← Software and systems

What the tool reads and under whose permissions. Where the data is stored and processed, what it keeps and for how long, and whether any of it trains models. Which other companies handle it, how each figure in an answer traces back to records, and how you get everything back and deleted when you leave. A vendor should answer all of that in writing, and the terms that matter belong in the contract. For a timber business, two of those questions deserve the most time. Permissions matter because the tool can only be as careful as the systems it reads. Traceability matters because a figure nobody can check can't be used.

What the usual answer says

Published checklists now cover most of the list above: storage, retention, training, subprocessors and exit. One on DEV Community even asks buyers to confirm that source-system permissions are "enforced per user at query time". Two things are missing from the results we read: evidence on how often buyers get these answers into contracts, and any question about whether the tool's answers can be traced to records.

Organisations in the 2026 benchmark study Say providers are transparent about data practices 81% Require contractual terms on ownership and liability 55% Question Answered in writing In the contract What the tool reads, under whose permissions Where data is stored and processed What it keeps, and for how long Whether any of it trains models Which other companies handle it How each figure traces back to records How you get it back and deleted on exit
Most organisations trust their AI providers' answers; far fewer put them in the contract. Each question below needs a written answer, and the terms that matter belong in the contract. Diagram: Quarri.

Buyers trust more than they contract

Cisco sells security and data-governance products. In September 2025 it surveyed over 5,200 IT, technology and security professionals with data privacy responsibilities, across 12 markets, for its 2026 Data and Privacy Benchmark Study. It found 81% of organisations say their generative AI providers are transparent about data practices. "However, only 55% require contractual terms that define data ownership, responsibility, and liability". At the same time, 70% "acknowledge risk exposure from the use of proprietary or customer data in AI training". The study's reading, framed by regulation, is that "informal assurances are no longer enough".

The sample is privacy staff in organisations large enough to employ them, and the study gives no figure by company size. A timber business of a few hundred people also has less bargaining power: large vendors rarely change their paper for it. That makes the question it controls directly, what gets connected and under whose permissions, the more important one.

Permissions

Microsoft's privacy documentation for its workplace AI assistant, updated in August 2026, is a good example of a written answer. It says the assistant "only surfaces organizational data to which individual users have at least view permissions". The same page adds: "It's important that you're using the permission models available in Microsoft 365 services, such as SharePoint, to help ensure the right users or groups have the right access to the right content within your organization".

That second sentence is the catch. A tool that respects permissions shows each user whatever the source systems already let them see. In our experience, a shared finance login or a drive of contracts open to the whole office is common in mid-size businesses. A chat window makes that exposure far easier to reach. So ask the vendor how permissions are enforced, then check your own. Before connecting a system, list what each role can see in it, and connect the narrowest scope first.

Traceability

Ask whether every figure in an answer links to the records it came from, and whether the user can open them. For stumpage paid, volumes or margins, a figure without its records has to be recomputed before anyone acts on it, which takes back the time the tool was meant to save. A vendor that can't show this in a demo on your data is unlikely to add it after signing.

The rest, in writing

Ask where the data, copies and processing sit. Ask how long prompts, answers and logs are kept, and whether you can set that period, since questions often contain the sensitive part. Ask whether anything trains or improves models, and what the default is. Microsoft's page, for example, says "Prompts, responses, and data accessed through Microsoft Graph aren't used to train foundation LLMs". Ask which other companies handle the data, including the model provider, and how you will hear of changes. And ask in what format the data, definitions and rules come back on exit, and when deletion is confirmed. Our piece on whether AI trains on company data covers training in more depth.

When it doesn't apply

AI tools used only on public information, such as market reports or published prices, carry less risk, though retention still matters if staff paste in anything private. Very large vendors may not negotiate terms, in which case the published terms are the contract and should be read as such.

Quarri's security page sets out how customer data is isolated, encrypted and audited.

Sources

  1. Cisco, "A Shifting Paradigm: Governance in the Age of AI. Cisco 2026 Data and Privacy Benchmark Study", survey September 2025: cisco.com
  2. DEV Community, "Before You Sign: How to Audit an AI Vendor's Data Practices": dev.to
  3. Microsoft Learn, "Data, privacy, and security for Microsoft 365 Copilot", updated 18 August 2026: learn.microsoft.com

Quarri is an AI-native data platform for the timber supply chain. It connects buying, production, sales and inventory for forest management, sawmill, wood products and pulp, paper and packaging operators.

See it on your own data.

Live in two weeks, on the systems you already run.