Data practice · 28 Sep 2026

Does AI train on your company's data?

← Data practice

It depends on the plan and the account, not the product. The large providers' business and API terms we checked forbid training their models on what customers send and receive. Their consumer plans may allow it, depending on the user's settings, and may keep the data longer when they do. So for a timber business the answer turns on which accounts its people use. A company can sign terms that forbid training and still have stumpage prices, margins and customer names in personal accounts whose owners said yes.

What the usual answer says

The usual answer is reassuring: enterprise AI providers don't train on customer data by default, so check the terms, opt out where possible and use enterprise versions with a data protection agreement. That is right for the accounts the company provides. It says nothing about the accounts it doesn't. And not every lawyer treats training as simply bad: Greenberg Traurig argued in January 2026 that "There are myriad instances when both the customer and the vendor might benefit from allowing models to train on customer data", such as security tools. That is a contract choice made on purpose, which is different from data slipping into a personal account.

Business and API plans under the commercial terms Personal plans Free, Pro and Max Training Not on customer content inputs or outputs Retention Agree a period in the contract alongside no training on your data Training If the user allows it a choice made per account Retention Five years if allowed 30 days if not Covers the accounts the company provides. The risk: personal accounts used for work.
The same assistant has two sets of terms. The company contract covers the accounts it provides, not personal accounts used for work. Anthropic's terms as published in 2025; terms change. Diagram: Quarri.

Two sets of terms for one assistant

One provider's published terms show the split clearly. Quarri builds on this provider's models, so we have an interest in how its terms read. Anthropic's Commercial Terms of Service, effective June 2025, state: "Anthropic may not train models on Customer Content from Services". Customer content there means both the inputs a customer sends and the outputs it receives.

Its consumer update of 28 August 2025 applies to "users on our Claude Free, Pro, and Max plans". It gives those users "the choice to allow their data to be used to improve Claude", and adds: "We are also extending data retention to five years, if you allow us to use your data for model training". Users who don't opt in stay on "our existing 30-day data retention period". The update states that it does "not apply to services under our Commercial Terms". These are the terms as published at the time of writing, and they change.

Microsoft's privacy documentation for its workplace assistant, updated in August 2026, states that "Prompts, responses, and data accessed through Microsoft Graph aren't used to train foundation LLMs". That is a commitment for the organisational product; check the consumer terms separately.

Where the risk sits

The risk is personal tools used for work. Microsoft and LinkedIn's 2024 Work Trend Index, from a survey of 31,000 people in 31 countries, found that "78% of AI users are bringing their own AI tools to work", and "it's even more common at small and medium-sized companies (80%)". Microsoft sells workplace AI, so it has an interest in that finding, and bringing your own tool doesn't always mean a personal account. But it is the gap a company contract doesn't cover.

Training is not the only question

Retention matters as much. The consumer update itself ties longer retention to training, and data that is never used for training can still be stored and reviewed under a provider's usage policies. And answering a question from a company's data is not training. A tool that looks up the company's records to answer a question (retrieval) uses the data at the moment of asking; the model itself doesn't learn from it. Training a private model on the company's own data is a third thing, a deliberate choice, with the model and data kept to the company.

What to do

Check training and retention per plan, in the provider's terms, not per product name. Provide business accounts for anyone who uses AI for work, so personal accounts aren't needed, and make them good enough that people use them. Say plainly which data may not go into personal tools: prices, margins, customer and supplier details, landowner information. Write no training on our data and a retention period into the contract, and review the approved tools when terms change.

When it doesn't apply

AI used only on public information, such as published market reports, carries little training risk. Tools that run entirely on the company's own hardware send nothing to a provider.

Quarri's security page sets out how customer data is isolated, encrypted and audited.

Sources

  1. Anthropic, "Commercial Terms of Service", effective 17 June 2025: anthropic.com
  2. Anthropic, "Updates to Consumer Terms and Privacy Policy", 28 August 2025: anthropic.com
  3. Microsoft Learn, "Data, Privacy, and Security for Microsoft Copilot", updated 18 August 2026: learn.microsoft.com
  4. Microsoft and LinkedIn, "AI at Work Is Here. Now Comes the Hard Part", 2024 Work Trend Index, 8 May 2024: microsoft.com
  5. Greenberg Traurig, "Is It Ever Okay for AI Providers to Train on Your Company's Data? Yes!", January 2026: gtlaw.com

Quarri is an AI-native data platform for the timber supply chain. It connects buying, production, sales and inventory for forest management, sawmill, wood products and pulp, paper and packaging operators.

See it on your own data.

Live in two weeks, on the systems you already run.