Research Note · AI · Negotiation

AI usage-based pricing negotiation.

The 2026 enterprise AI contract is per-token consumption, not per-seat. This note breaks the negotiation into the seven commercial levers that actually move realised cost — rate cards, volume tiers, caps and floors, commitment discounts, rollover, true-forward protection, and capacity reservation — and shows how the buyers who outperform treat AI as commodity commit procurement rather than software licensing.

By James Hill-WoodUpdated Sep 20229 min readAI research cluster
Bottom line

There is no single "cheapest" AI vendor — only a better-negotiated contract. The 2026 negotiation surface has moved from seat count to consumption commit, ramp profile, overage rules, and capacity reservation pricing. The buyers who outperform treat AI as commodity commit negotiation on the AWS EDP model, not software licensing on the Microsoft EA model — worth a median 24 points below opening commit.

01 Key findings

  1. AI spend has shifted from seat to token. Per-seat chat licences fell from roughly 72% of AI spend in 2024 to 38% in 2026; API token consumption rose from 18% to 42% and now frequently exceeds seat spend within 12 months of rollout.

  2. Capacity reservation is the single largest commercial decision. Reserved throughput (Azure PTU, Bedrock Provisioned Throughput, Anthropic Capacity Reservation, Vertex Provisioned Throughput) breaks even versus per-token at roughly 100–200 sustained tokens/sec, with realised discounts of 25–45% — but carries utilisation risk.

  3. Overage rate is the most frequently missed term. The vendor default is list rate above commit, which can double the effective rate. The achievable position is contracted rate above commit with an automatic upgrade above 125%.

  4. The ramp profile decides three-year TCO, not the unit price. Aggressive vendor ramps secure revenue growth; the achievable customer position is year one at current consumption plus 20–35%, year two at 1.8–2.2x, year three at 1.5–1.8x.

  5. Concurrency beats sequencing. Running at least two vendors in parallel for 60+ days creates the documented alternative pricing that vendor flexibility correlates with directly.

02 The seven pricing levers

Every 2026 AI consumption contract resolves into seven commercial levers. Each has a vendor default position and an achievable customer position; the gap between them is the negotiation.

LeverWhat it controlsVendor defaultAchievable position
Rate cardPer-million input/output token unit priceList rate, model-by-modelBlended committed rate, price-protected for term
Volume tiersDiscount stepped by cumulative spendCoarse tiers, list between stepsCustom tiers set to credible consumption
Commitment discountDiscount for annual dollar commit15–22% year one28–38% by year three on ramping commit
Caps & floorsCeiling on unit price; minimum spendNo cap; high spend floorPrice cap on renewal; floor at true consumption
RolloverTreatment of unused committed spendUse-it-or-lose-it, expires quarterlyAnnual rollover of underspend within term
True-forward protectionHow overspend resets the commitRetroactive true-up to listTrue-forward only; no retroactive repricing
Capacity reservationReserved throughput vs pay-as-you-goPeak-sized, full-term lockBaseline-sized, burst to PAYG on same model

03 Capacity reservation versus pay-as-you-go

The largest single decision is whether to commit to reserved capacity or stay on per-token pay-as-you-go. It turns on three numbers: sustained tokens per second across production workloads, the realised discount on reserved capacity, and the burst pattern the workload exhibits.

Reservation breaks even at roughly 100–200 sustained tokens/sec, with realised discounts of 25–45% at scale. The trade-off is utilisation risk: reserved capacity running at 30% utilisation is more expensive than per-token on the same workload. Size reservations against the steady-state baseline and handle burst with per-token billing on the same model.

The PTU sizing trap

Azure OpenAI PTUs are sold in throughput per minute. The standard mistake is to size against peak load, which over-provisions for the average workload — the single most common waste pattern in 2026 Azure OpenAI deployments, with a typical 30–60% of PTU capacity unused on a steady-state basis. Size against median sustained load and route peaks to pay-as-you-go. See our Azure MACC analysis for the commercial structure PTUs sit within.

04 Commit structure and ramp profile

Enterprise AI contracts above $500K annual commit increasingly mirror the AWS EDP structure: multi-year term, annual ramping commit, tiered discount on cumulative spend. The ramp profile is the most contested term — vendors want aggressive ramps to secure revenue, customers want conservative ramps because early-adoption uncertainty is real.

YearTypical commit patternDiscount tierAchievable ramp
Year 1$500K to $1M15 to 22%Current consumption + 20–35%
Year 2$1.2M to $2.5M22 to 30%1.8–2.2x of year one
Year 3$2M to $4.5M28 to 38%1.5–1.8x of year two

Negotiate the ramp, not just the price. The ramp determines three-year TCO more than the unit rate, and a credible year-one anchor keeps later-year commitments defensible if adoption lags.

05 Overage rules and the uncapped-usage trap

When consumption exceeds the contracted commit, the overage is billed at one of several rates. The default vendor position is list rate above commit; the achievable position is contracted rate above commit, with an automatic upgrade above 125% that also secures the higher commit for the vendor.

Overage structureCustomer impactVendor default?
List rate above commitPunitive, can double effective rateYes (initial position)
Contracted rate above commitPredictable, preferredNegotiable
Tiered (contracted to 125%, list above)Acceptable with monitoringAcceptable
Automatic commit upgrade above 125%Best long-term economicsAchievable
Soft cap with notification, no auto-billBest for cost governanceHard to achieve
The uncapped-usage trap

Uncapped per-token consumption with list-rate overage is the fastest way to lose an AI budget. A workload that grows faster than the commit is silently billed at list above the committed volume, erasing the negotiated discount exactly when spend is highest. Insist on contracted-rate overage and a soft cap that notifies before it auto-bills — the single most frequently missed term, and the one that produces the most painful overage invoices.

06 Data rights and exit terms

The contractual question legal teams care about most is the training opt-out. The frontier vendors (Anthropic, OpenAI, Microsoft, Google) all default to no-training on Enterprise-tier customer data, but the mechanism varies: Anthropic and OpenAI bake it into the standard Enterprise MSA; Microsoft inherits the M365 DPA position; Google inherits the Workspace DPA position. Negotiate explicit retention windows (typically 30 days for abuse monitoring, with a no-retention option), explicit audit rights, and explicit data-deletion terms at exit.

Exit is harder for AI than for SaaS. Custom GPTs, Projects, prompt libraries, agent definitions and tool integrations do not port between vendors. The operational mitigation is architectural: keep prompt logic and RAG data in customer-controlled storage and call the model as a stateless service. The exit terms that materially matter are notice period for non-renewal (90 days typical, push to 60), data deletion at exit (30 days post-termination is achievable), export assistance (best-efforts achievable, explicit SLA harder), and the right to deactivate auto-renewal mid-term (achievable on Enterprise tier).

07 Negotiation framework

Four moves consistently deliver across advisor-led deals on OpenAI, Anthropic, Microsoft, Google and AWS Bedrock contracts. Weight them to your situation before signing.

Move 01

Run vendors in parallel

Vendor pricing flexibility correlates directly with documented alternative pricing. Pilot at least two vendors for 60+ days before signing either — the leverage evaporates once you have committed.

Move 02

Separate seat from API commit

They are different commercial constructs with different economics. Negotiating them together obscures the per-seat versus per-token picture and lets the vendor cross-subsidise the weaker line.

Move 03

Route through existing commit

Claude on Bedrock burns AWS EDP; Azure OpenAI burns Azure MACC; Vertex burns Google Cloud commit. Token economics are identical but the commercial accounting is materially different.

Move 04

Time the close to quarter end

Microsoft FY ends 30 June, AWS Q4 ends 31 December, and OpenAI and Anthropic follow calendar quarters. The last two weeks of a fiscal quarter deliver consistent pricing flexibility.

08 Our recommendation

Reserve capacity
When load is steady

You run sustained production workloads above the break-even throughput. Size reservations against the steady-state baseline, burst to pay-as-you-go on the same model, and avoid the peak-sizing over-provisioning trap.

Stay PAYG
When usage is volatile

You are early in rollout with unpredictable consumption. Keep per-token billing, negotiate contracted-rate overage, and add capacity reservation later once the workload pattern stabilises.

Ramping commit
When scaling fast

You have a credible growth trajectory above $500K. Anchor year one at current consumption plus 20–35%, protect the unit rate with a cap, and secure rollover on any underspend.

09 Negotiation sequencing

The single highest-value process choice for AI consumption buyers:

Concurrent Recommended

Pilot at least two vendors in parallel for 60+ days, with each aware a primary decision is live. This produces the documented alternative pricing that vendor flexibility correlates with — and the competitive tension that drives best-in-class terms.

Sequential Weaker

Sign one vendor, then benchmark. Each later quote can undercut the last, but pressure on the incumbent collapses once you have committed, and total leverage falls sharply.

Treat AI procurement as commodity commit negotiation

Our AI procurement practice coordinates timing, benchmarking and strategy across every frontier vendor at once.

Request AI advisory →

The Licensing Edge

Weekly AI and licensing intelligence for enterprise IT leaders. 3,000+ subscribers.