Someone on your team is about to paste your customer list into a chat window. Or connect your CRM to an AI note-taker. Or upload three years of support tickets so a model can "learn your tone." The security review asked whether the vendor is SOC 2 compliant. Nobody asked the question that actually decides your exposure: what rights did you just grant, and to how many companies?
Encryption protects data from people who were never supposed to have it. A contract governs what happens to data you handed over on purpose. Those are different risks, and only one of them shows up in a security questionnaire.
The expensive clause is the one you already agreed to
Most AI tools are bought the way software has always been bought: someone needs a result this quarter, a free tier works, a card gets entered, and the terms are accepted by clicking. The terms are usually reasonable. But "usually reasonable" is a bet, and the size of the bet scales with how much of your business flows through the tool.
Here is the practical version of the risk. If a vendor holds broad rights to use your inputs for model improvement, then your pricing logic, your client names, your internal playbooks, and your unreleased positioning all become raw material for a system you don't control. You will not get a notification. You will not see a diff. The rights were granted at signup.
Regulators have been explicit that quietly widening those rights after the fact is a legal problem. In a February 2024 post, the FTC warned that adopting more permissive data practices and informing users only through a retroactive amendment to the terms of service or privacy policy may be unfair or deceptive. That's a useful signal about vendor behavior, but it isn't a shield for you. Enforcement runs on the regulator's timeline, not your incident timeline.
The consumer-versus-business gap is where most teams get caught
The same brand can run two completely different data policies depending on which door you walked through. Business and API tiers tend to carry no-training commitments. Consumer tiers often don't, or they default to training unless someone toggles a setting.
Read the providers directly rather than trusting summaries, including this one:
- OpenAI documents that data sent to its API is not used to train its models unless you explicitly opt in, with separate handling for abuse-monitoring logs. Its enterprise privacy page covers business tiers.
- Anthropic's commercial terms govern API and business use, and are distinct from the consumer policies that apply to individual Claude accounts.
Now apply that to how your company actually works. If your ops lead runs analysis in a personal account on a personal laptop because it's faster than waiting for procurement, your enterprise agreement does not cover that session. The contract you negotiated protects the traffic that goes through the account you negotiated it for. Everything else runs on whatever terms that individual accepted.
This is a policy and enablement problem before it's a legal one. Deciding which tier each team uses, and training people on why it matters, is cheaper than discovering the gap later. If you're standing that up from scratch, our learning resources and the operator conversations in our community are a reasonable place to start.
Retention is not deletion, and deletion is not erasure
Three separate clocks run on your data, and vendors often describe only the first one.
- Active storage. How long your prompts, files, and outputs sit in the vendor's systems during normal use.
- Log and safety retention. A separate window, often shorter or longer than the first, covering abuse monitoring and debugging. This one survives your "delete conversation" click.
- Derived artifacts. Embeddings, vector indexes, fine-tuned weights, evaluation sets, cached retrievals. These are built from your data but are not your data, and generic deletion language frequently misses them.
Point three is the one worth arguing about. The FTC has treated models themselves as forfeitable: in its Everalbum matter, the order required the company to delete the models and algorithms it developed using photos and videos uploaded by its users, not just the underlying images. Regulators understand that the model is the finished good. Your termination clause should reflect the same understanding: ask explicitly whether deletion covers fine-tunes and embeddings derived from your content, and how the vendor evidences it.
If a vendor can't answer that in writing, treat it as a finding, not a formality.
You are not buying from one company
Every AI vendor is a stack. There's a model provider underneath, a cloud host underneath that, plus vector storage, observability tooling, and often a transcription or OCR service you've never heard of. Each one is a subprocessor, and each one is a place your data lives.
Under Article 28 of the GDPR, processors need authorization before engaging another processor and must flow equivalent obligations down the chain. Even if the GDPR doesn't apply to you, that structure is the right mental model. California's regime imposes comparable contractual requirements on service providers, and the California Privacy Protection Agency publishes the current rules.
Two questions cut through the ambiguity: Where is the current subprocessor list published, and how are we notified before it changes? A vendor that maintains a public, versioned list has thought about this. A vendor that emails you a PDF has not.
This matters more as you automate. A single tool holds one copy of your data. A chain of connected automated workflows pushes the same records through five vendors and twenty subprocessors, and each hop inherits a different contract.
Read the agreement in this order
Skip the preamble. Go straight to these, in this sequence, because each answer changes what the next one means:
- Ownership and license grant. Not "who owns the data," which is almost always you. The real question is the scope of the license you granted: is it limited to providing the service, or does it extend to improving, developing, or training?
- Training rights and opt-out mechanics. Is the no-training position a contractual commitment or a setting someone can flip? A default you must configure is weaker than a term the vendor is bound by.
- Output ownership and indemnity. Do you own generated outputs, and will the vendor defend you against third-party IP claims arising from them? Read the carve-outs; they're where the protection actually lives.
- Retention, deletion, and derived artifacts. All three clocks above, plus what happens on termination and how deletion is certified.
- Subprocessors and change notice. Published list, notification window, and whether you can object.
- Amendment mechanics. Can the vendor change material data terms unilaterally with notice-by-webpage? This clause quietly governs every clause above it.
- Incident notification and audit. Notification windows measured in hours or days, and what evidence you can request without a lawsuit.
Bring this to counsel rather than instead of counsel. Nothing here is legal advice, and the right redline depends on your jurisdiction, your data, and your leverage.
What to do when you have no leverage
A ten-person company is not renegotiating terms with a major platform. Fine. Controls you own are often stronger than clauses you can't get:
- Classify before you connect. Decide which data tiers are allowed into which tools. Most teams find that the highest-value use cases don't require the most sensitive data.
- Strip identifiers at the boundary. Replace names, account numbers, and contact details before the request leaves your systems. A model doesn't need to know who the customer is to summarize the ticket.
- Buy the tier that carries the terms you need. Zero-retention and no-training options frequently sit behind a paid plan. Compare that cost against the cost of the exposure, not against free.
- Own the layer that holds the sensitive parts. Sometimes the correct answer is a system you control, with the model doing only the narrow task that genuinely needs it. That's a build decision, and it's the reason some teams end up with purpose-built internal software instead of a stack of subscriptions.
For structuring this across a real vendor portfolio, that's the kind of work our AI consulting engagements handle, and we cover the practical version of it at Tech Lab Miami events.
Make it a repeatable process, not a one-time panic
Inventory every AI tool touching company data, including the ones nobody expensed. For each, record the tier, the training position, the retention windows, the subprocessor list, and the termination-deletion language. Review it on a schedule, because terms change and your inventory grows faster than your memory of it.
If you want an external structure for this, the NIST AI Risk Management Framework explicitly addresses risks from third-party software, data, and AI supply chains, and its Generative AI Profile extends that to generative systems. Both are voluntary and free. Adopting even a light version gives you something defensible to show a client, an insurer, or an acquirer.
The point isn't to slow adoption down. It's that the teams moving fastest with AI a year from now will be the ones who knew what they signed.
Frequently Asked Questions
Can I negotiate the terms of my AI vendor contract?
Often, yes, and more often than teams assume. Data retention windows, training rights, subprocessor notification periods, and deletion certification are all commonly negotiated, especially with mid-market vendors competing for your business. Large platforms are less flexible on standard terms but frequently offer stricter configurations on higher tiers instead. If you can't move the paper, move the plan. Either way, have counsel review before signing.
What should I do if I think a vendor is misusing my data?
Start with the documents: the agreement in force at signature, any amendments since, and the version of the privacy policy that applied when the data was submitted. Preserve them, because vendors update pages without archiving them. Then check your audit and incident-notification rights to see what you can formally request. Bring counsel in early, particularly if regulated or personal data is involved, since your notification obligations may be running independently of the vendor's.
How can I protect my data when working with an AI vendor?
Layer it. Get the contractual terms right on training, retention, deletion, and subprocessors. Then control what actually leaves your systems through data classification and identifier stripping, so the strongest terms are protecting the smallest amount of sensitive material. Then verify: keep a live tool inventory, re-check terms on a schedule, and confirm that deletion requests were honored rather than assuming it. Controls you operate are the ones you can prove.
Does a no-training commitment mean my data isn't stored at all?
No, and conflating the two is a common mistake. A vendor can commit to never training on your inputs while still retaining them for abuse monitoring, debugging, or legal hold. Those are separate clauses with separate timelines. If you need storage minimized, look specifically for zero-retention terms and confirm what exceptions survive them.
Visuals

