Research Note · AI · Contracts

AI IP ownership: who owns what you create with AI.

Enterprises are embedding generative AI into revenue products, yet the contracts rarely say who owns the output. This note dissects copyright eligibility, the three vendors' output and training positions, indemnity scope, and the clause language that keeps your organisation's IP defensible.

By James Hill-WoodUpdated Jan 202410 min readAI research cluster
Bottom line

Owning AI output is not the same as being able to defend it. Copyright attaches only where a human contributes real creative direction; vendor terms diverge sharply on whether they train on your inputs; and indemnities are riddled with carve-outs. Treat output ownership and data isolation as non-negotiable, and force fine-tuned model ownership and broad-scope indemnity into every material AI contract.

01 Key findings

  1. Ownership and copyright are different rights. A vendor can assign you "the output" while that output remains uncopyrightable — unprotectable against scraping or competitors — because it lacks human authorship.

  2. Training rights are the quiet risk. Unless you secure data isolation, proprietary inputs and outputs can flow into a vendor's training corpus and inform models your competitors then use.

  3. The three vendors are not interchangeable. OpenAI is the most customer-friendly on ownership and indemnity; Google retains broad training rights by default; Microsoft assigns output but carves out open-source material.

  4. Fine-tuned models are underprotected. You may invest millions adapting a model, then find the vendor retains rights to the process and outputs — eroding the advantage you paid for.

  5. Every indemnity has limits. Caps, modification exclusions and open-source carve-outs mean a headline indemnity can leave your true exposure uncovered.

03 Vendor output positions

The three largest vendors take markedly different positions on output ownership and training rights, each with distinct commercial consequences. Match the vendor posture to how sensitive your inputs are and whether you can pay for isolation.

VendorOutput ownershipTraining on your dataCopyright indemnityKey limitation
OpenAIAssigned to customerNot used unless opt-inBroad; covers core outputsExcludes materially modified outputs; not embeddings
GoogleCommercial licence to useRetained by default; isolation at higher tierSpecific service tiers onlyData isolation typically needs >$500k annual spend
MicrosoftAssigned to customerNot used without opt-inCopilot Copyright CommitmentOpen-source training material carved out
Read the plain-language version

OpenAI: you own outputs, your data is isolated, we defend copyright claims. Google: you have a licence to use outputs, and we may train on your data unless you pay extra. Microsoft: you own outputs and we defend copyright claims — except where a Copilot output reproduces open-source material from the training corpus.

04 Fine-tuned model ownership

Output ownership is only half the story. The second battleground is the fine-tuned model. When you adapt a vendor model with proprietary training data, custom examples and domain language, who owns the resulting weights and parameters? The answer is vendor-dependent and often contractually unclear.

OpenAI's fine-tuning service grants customers ownership of the fine-tuned weights: you deploy it, run it in production and restrict access. Google Vertex AI and Microsoft Azure OpenAI Service are less transparent — weights are typically owned for deployment, but the vendor retains rights to use the fine-tuning process and outputs to improve its own services. If your model embodies proprietary customer, manufacturing or competitive-intelligence data, vendor-retained rights quietly erode the advantage you invested in.

Contamination risk

Feed proprietary information into a training-enabled system and it can inform future models that competitors access. In 2024, organisations found proprietary code and business logic fed into enterprise Copilot influenced model behaviour in ways that leaked organisational patterns. Demand no-training commitments, regional data residency, audit rights, and permanent deletion.

05 Indemnity scope

A copyright indemnity is the vendor's commitment to defend and indemnify you against third-party infringement claims. It matters here because training data includes vast copyrighted material, and outputs can resemble or reproduce it. But scope varies dramatically — and no indemnity is unconditional.

Indemnity tierWhat it coversTypical exclusionsWho offers it
BroadAll copyright claims related to your outputsCustomer modifications; customer-supplied training dataOpenAI
NarrowOnly where output matches vendor-provided training dataOpen-source, user-provided or third-party materialGoogle (specific tiers), Microsoft (open-source carve-out)
NoneNo commitment; customer bears all riskEverythingMany smaller vendors
Watch the cap

Every indemnity contains caps on liability, exclusions for customer modifications, and reasonable-security requirements. An indemnity capped at $10M against $100M of exposure provides false comfort — model the cap against your realistic worst case, not the headline promise.

06 Contract traps

The recurring failures we see in enterprise AI agreements are less about missing clauses than about clauses that read well but do nothing. Watch for these.

Trap 01 · "Output ownership" without copyright

Vendors assign "output" while staying silent on copyright, commercial use rights, or training rights. You can own something you cannot defend. Insist the assignment names all three.

Trap 02 · Default training rights

Silence on training usually means the vendor retains the right to use your inputs. Isolation is an add-on, often gated behind a spend threshold. Make no-training the default, in writing.

Trap 03 · The open-source carve-out

An indemnity that excludes open-source training material leaves code-generation outputs exposed to the very claims most likely to arise. Price that gap before you rely on the commitment.

Trap 04 · Illusory caps

Liability caps far below your exposure, or aggregate caps across all claims, convert an indemnity into theatre. Push for per-claim treatment and caps sized to real risk.

07 Negotiation checklist

Five clauses turn a standard AI contract into a defensible one. Work them in this order; the first two are where sensitive deployments live or die.

Clause 01

Exclusive output assignment

Vendor assigns all right, title and interest in output to the customer, who may use, modify, publish and commercialise it — regardless of later modification or integration.

Clause 02

Data isolation & training ban

Vendor shall not use input, output or derivatives to train, improve or develop any model, feature or service without explicit written consent — extended to affiliates and partners.

Clause 03

Broad-scope indemnity

Vendor defends and indemnifies against third-party copyright, patent or trade-secret claims on output, with a stated cap and no aggregate cap across claims.

Clause 04

Fine-tuned model ownership

Any model, weights or derivative created by fine-tuning on customer data is owned exclusively by the customer; vendor retains only support-and-maintenance access.

Clause 05

Deletion & audit rights

Vendor deletes all customer data within 30 days of request or termination, certifies deletion in writing, and grants quarterly audit access to verify isolation and compliance.

08 Our recommendation

Choose OpenAI
When protection wins

You want the current gold standard — output assignment, training isolation and broad indemnity in one package. Confirm the indemnity's core-output scope and note the modification exclusion before relying on it.

Choose Google
When you can pay for isolation

You have the scale to reach the data-isolation tier and the discipline to negotiate it. Without that tier, assume your inputs train Google's models — unacceptable for sensitive IP.

Choose Microsoft
When Copilot fits the estate

You value output assignment plus the Copilot Copyright Commitment across users and enterprise. Price the open-source carve-out carefully if you generate code at scale.

09 Negotiation priorities

Not every buyer can move a hyperscaler off standard terms. Sequence your asks so the non-negotiables come first and the tail-risk protections follow.

Tier 1 Non-negotiable

Output ownership assignment and data isolation. Without these, the contract is unacceptable for any sensitive use case — walk before you sign around them.

Tiers 2–3 Push hard

Copyright indemnity with adequate caps and deletion rights (now table-stakes at enterprise scale), then fine-tuned model ownership, audit rights and no-aggregation caps. Enterprise tiers cost 30–50% more — justified by the IP you protect.

Strengthen your AI contracts

Our AI Procurement Advisory practice reviews vendor terms, identifies IP risk, and leads negotiation to secure output ownership, isolation and indemnity.

Explore AI procurement advisory →

The Licensing Edge

Weekly AI and licensing intelligence for enterprise IT leaders. 3,000+ subscribers.