Cloud SLA negotiation: what enterprise buyers must demand.
Cloud service level agreements are written by vendors for vendors: exclusions that void coverage, credit percentages unrelated to business impact, and dependencies few buyers negotiate. This note sets out the SLA terms to demand — uptime, service-credit structure, exclusions, remedy caps, chronic-failure termination, and how availability is actually measured.
Standard cloud SLAs are asymmetric by design. 99.9% uptime permits 8.7 hours of downtime a year, and the 10–30% service credit it triggers rarely approaches real business loss — an outage that costs millions may return credits of $5,000–10,000. The terms are negotiable. Buyers who demand tiered credits, narrowed exclusions, incident-duration measurement and chronic-failure termination — before signing — secure protection the default contract never provides.
01 Key findings
Uptime percentages are widely misread. 99.9% permits 8.7 hours of downtime per year (43 minutes per month); 99.99% permits 52 minutes per year; 99.999% permits 5.2 minutes. The gap between tiers is commercially material and worth explicit negotiation for critical workloads.
Standard credits are not remedies. Credits of 10–30% of monthly fees bear no relationship to business loss. Mission-critical workloads need graduated credits — 50%, 75% or 100% for extended outages — plus remedies beyond credits.
Exclusions are where the risk lives. Planned maintenance, DNS failures, force majeure and the broad “customer error” carve-out routinely void coverage. The maintenance exclusion is the most damaging: providers may patch during your peak hours with no recourse.
Dependencies compound. Five services each at 99.9% combine to roughly 99.5% (0.9995). Most large applications run below the availability the buyer believes it purchased.
Monthly measurement creates perverse incentives. Because credits trigger on cumulative monthly thresholds, providers optimise to keep any single incident just short of the trigger rather than to prevent recurrence. Shift measurement to incident-duration thresholds.
02 What SLAs really guarantee
A cloud SLA specifies an uptime percentage, what happens if it is missed (typically service credits), and the events excluded from the guarantee. Crucially, it is defined in terms of infrastructure uptime, not business outcome. “99.9% uptime” does not mean your application is available 99.9% of the time — it means the provider’s compute, storage and networking components meet a technical benchmark 99.9% of the time. Those are not the same thing.
Standard SLAs ignore application-level availability, third-party dependencies and your own configuration. Misconfigure a security group and block inbound traffic, and the provider is still meeting its SLA. Deploy a single point of failure across zones, and the SLA remains satisfied. Most enterprise incidents stem from architecture, configuration or excluded dependencies — not from provider infrastructure failure. Understanding this asymmetry is the foundation of intelligent SLA negotiation.
03 SLA terms to demand
Beyond the headline uptime figure, five terms decide whether an SLA actually protects the business. Demand each explicitly, and specify the standard rather than accepting the provider default.
| Term | Provider default | What to demand | Why it matters |
|---|---|---|---|
| Uptime tier | 99.9% (8.7 hrs/yr) | 99.99% or 99.999% for critical workloads | Each tier cuts permitted downtime roughly 10× |
| Service credits | 10–30% of monthly fees | Graduated 50/75/100% by outage severity | Default credit rarely approaches business loss |
| Scope | Core “managed” services only | All deployed services + recommended HA configs | Narrow scope voids coverage on your real estate |
| Exclusions | Broad: maintenance, DNS, “customer error” | Only truly uncontrollable events | Exclusions are the primary route to zero recourse |
| Dependency SLAs | Per-service, buyer’s problem | Composite SLA or enhanced per-component terms | Combined availability is lower than any single part |
04 The toothless-credit trap
Standard credits range from 10% to 30% of monthly service fees for the affected service. If a service costing $100,000 a month is unavailable for four hours — breaching the 99.9% SLA — the buyer receives $10,000–30,000 in credits. That is almost never equivalent to the loss. The structure of the calculation matters as much as the percentage: standard SLAs award flat credits by availability band, so a five-minute outage and a one-hour outage can return the same amount.
A credit is a partial refund, not compensation. A financial-services firm losing its trading platform for one hour may lose millions; the provider’s standard credit totals perhaps $5,000–10,000. For mission-critical workloads, demand: (1) a higher base credit (minimum 50%), (2) escalating credits — 75% for outages of 30+ minutes, 100% for 1+ hour, (3) remedies beyond credits (cash, extended term, or termination rights), and (4) mandatory root-cause analysis for any incident exceeding 30 minutes.
05 Dangerous exclusions
Exclusions are where the real risk sits. Standard SLAs exclude planned maintenance, DNS failures, force majeure, external attacks, third-party service failures, and anything attributable to “customer” actions. Two exclusions do the most damage.
The planned-maintenance exclusion lets providers patch, roll updates and move capacity — continuously — entirely outside the guarantee. If maintenance lands during your peak hours, you have no recourse, and the SLA often does not even define what “maintenance” means. Demand: maintenance only outside declared business hours, at least 14 days’ notice, no more than one window per service per month during business hours, and customer approval of proposed windows.
The “user error” carve-out excludes issues from “customer configuration,” “customer application code” or “customer API usage” — language broad enough to exclude almost anything not caused by a provider code bug. If unclear documentation leads you to misconfigure, is that user error or provider error? Narrow the exclusion to apply only to customer application code and data, not to platform configuration the provider should document clearly.
06 How uptime is measured
Measurement mechanics quietly decide who pays. 99.9% uptime does not mean the service is down for 8.7 hours spread randomly — it means cumulative unavailability stays under the monthly threshold. Eight consecutive hours down in one month breaches it and triggers credits; 43 minutes down does not. Multi-region deployments compound the problem: most SLAs apply per region, so a failover target without negotiated terms carries no guarantee at all.
Monthly thresholds reward speed over prevention. Resolve an issue in 42 minutes and the provider avoids credits; take 44 and they apply — encouraging rapid mitigation at the expense of accurate investigation and recurrence prevention. Shift measurement to meaningful incident-duration thresholds: no incident should exceed 15 minutes; 30+ minutes triggers escalated investigation; 1+ hour triggers a joint incident review with the customer.
07 Negotiation checklist
Work these factors before signing. Weight them to workload criticality, and raise enhanced-SLA requirements at the start of negotiations — not as an afterthought.
Set the uptime tier deliberately
Map each workload to a tier and price the gap. AWS, Azure and Google Cloud can deliver 99.99% or 99.999% for accounts structured for high availability and credibly committed to multi-year spend.
Restructure credits and remedies
Replace flat bands with graduated credits, add remedies beyond credits, and secure termination rights if uptime falls below threshold for consecutive months.
Narrow every exclusion
Constrain maintenance to out-of-hours with notice and frequency caps, and limit “user error” to your code and data, not platform configuration.
Model the dependency stack
Calculate composite availability across every service in the architecture, then negotiate enhanced per-component SLAs or a composite guarantee — and cover every region you deploy to.
08 Our recommendation
Accept 99.9% but reject the default credit structure. Push for graduated credits and a narrowed maintenance exclusion — the two cheapest wins that materially reduce exposure.
Move to 99.99%+ with architect-validated criticality, escalating credits to 100%, root-cause obligations, and chronic-failure termination rights. Raise it at renewal, when providers hold pricing authority.
Model the dependency product, then either enhance each component SLA or secure a composite guarantee. Extend coverage to every region, including failover targets.
Close the gaps in your cloud SLAs
Our Cloud & FinOps practice reviews SLA structure, models credit exposure and negotiates enforceable enhancements across AWS, Azure and Google Cloud.
09 Frequently asked questions
What does 99.9% uptime actually mean in terms of downtime? 99.9% uptime permits 8.7 hours of downtime per year, 43 minutes per month, or approximately 8.6 seconds per day. 99.99% permits 52 minutes per year; 99.999% permits just 5.2 minutes. Many buyers mistakenly read 99.9% as near-perfect, but it permits over 8 hours of unplanned downtime annually — and for mission-critical systems the gap to 99.99% is worth negotiating explicitly.
Are SLA credits the same as actual financial remedies? No. Standard credits are typically 10–30% of monthly service fees, almost never equivalent to the business loss. A firm losing trading systems for an hour may lose millions while credits total $5,000–10,000. Negotiate higher percentages (50–100% for extended outages), remedies beyond credits, and termination rights if uptime falls below threshold for consecutive months.
What are the most dangerous SLA exclusions? Planned maintenance (often excluded entirely, even during business hours), DNS failures and route issues, user error and misconfiguration, force majeure and external attacks, and third-party dependencies. Multi-tier architectures often breach thresholds because of excluded dependencies rather than provider failure. Demand exclusion of only truly uncontrollable events, prior notice of out-of-hours maintenance, and coverage or explicit acknowledgment for dependent services.
Can we negotiate higher SLAs for strategic cloud workloads? Yes. AWS offers enhanced SLAs up to 99.99% for specific services; Azure offers Premium SLAs above standard tiers; Google Cloud offers custom SLAs for large accounts. Enhancement typically requires a multi-year commitment, architect-level validation of criticality, willingness to implement recommended resilience patterns, and a credible threat to move workloads — most accessible at major renewals.
Who should help me negotiate cloud SLAs? The work spans architecture, commercial terms and risk. Internal teams should include infrastructure leads, finance, procurement and legal counsel. Externally, Atonement Licensing advises exclusively on the buyer side, with former AWS and Azure commercial team members. Avoid advisors holding provider certifications or reseller relationships, which create conflicts of interest.
The Licensing Edge
Weekly cloud and licensing intelligence for enterprise IT leaders. 3,000+ subscribers.