COE 558Lecture 02Part 10
Service-level agreements and availability
Defines availability, maps SLA nines to allowed downtime, and shows how serial services combine into a compound SLA.
- Concepts
- 4
- Slides
- 63-68
- Reading
- 24 min
Why this part matters
The earlier parts of this lecture asked where a service should run and who should own the hardware. This last part asks a sharper question: how often will that service actually be there when a user calls it? Every Cloud service provider (CSP) sells the answer as a single number with a row of nines, and every architecture you design chains several such numbers together.
Turning nines into minutes of downtime, and composing them correctly across a pipeline, is a calculation you should expect on the exam. You will also need it the moment your research system spans an edge site, a fog node and a cloud region: each hop is another dependency, and each dependency spends part of your reliability budget.
By the end you can
- Compute availability from uptime and downtime, from successful and valid requests, or from MTBF and MTTR.
- Convert any number of nines into allowed downtime per day, week, month and year, and state the period convention you used.
- Compute the compound SLA of services in series and contrast it with the availability of redundant replicas in parallel.
- Distinguish SLI, SLO and SLA, and read the commitments and service credits of a real provider SLA.
Take a campus course-registration API and watch it for a 30-day month, which is 720 h. During that month it is unreachable twice, for a total of 2 h. For the other 718 h it answered requests. The fraction of the month it could answer is 718 / 720 = 0.9972, or 99.72%. That fraction is the service's Availability.
The general rule only names the two pieces. Uptime is the time the application can serve requests, and Downtime is the time it cannot. Availability is the share of uptime in the total:
The denominator is simply the length of the observation period, . Rearranging gives the form you will use most often, because it turns a promised availability into an allowance of failure: . For the registration API, (1 − 0.9972) × 720 h ≈ 2 h, which is exactly where we started.
Counting requests instead of hours
Large providers rarely have a clean on or off switch. A service may be up for most users while one shard fails. So they often measure availability by requests rather than by time: successful requests divided by valid requests. Google's SRE book gives a concrete case. A system that serves 2.5 million requests a day with a daily target of 99.99% may fail up to 250 of them. AWS combines both views: it measures success in windows of 1 or 5 minutes and averages the windows into a monthly uptime percentage.
Estimating availability before you have measurements
At design time there is no uptime log yet. Instead you can estimate availability from two failure statistics: the mean time between failures (MTBF) and the mean time to recover (MTTR).
Worked example
A component that fails every 150 days
Put both statistics in one unit
MTBF is 150 d = 3600 h. MTTR is 1 h.Apply the estimate
A = 3600 / (3600 + 1) = 3600 / 3601 ≈ 0.99972.Result
About 99.97%. The formula shows the two levers you have: fail less often (raise MTBF) or recover faster (cut MTTR). Halving the recovery time to 30 min buys as much as doubling the time between failures.
Recall
A component has an MTBF of 30 days and an MTTR of 2 hours. What is its availability, and which two levers raise it?
Start with the most common promise in cloud marketing, 99.9%, often called “three nines”. Its unavailability is 1 − 0.999 = 0.001. Multiply that by a period and you get the Downtime it allows: 0.001 × 8760 h = 8.76 h a year, 0.001 × 720 h = 43.2 min in a 30-day month, and 0.001 × 168 h ≈ 10.1 min a week.
Now add one more nine. An Availability of 99.99% has an unavailability of 0.0001, ten times smaller, so every downtime figure shrinks tenfold: 52.56 min a year, 4.32 min a month. That is the whole rule. Google's SRE book calls each additional nine “an order of magnitude improvement”, and the cost of getting there typically rises much faster than tenfold.
The table below repeats the calculation for each common tier. Like the lecture slide, it assumes a 365-day year, a 30-day month and a 7-day week. Every cell on the slide recomputes correctly with those conventions, and the figures agree with the availability table in the Google SRE book. The last column gives the example workloads AWS associates with each tier, so you can feel what a nine buys.
| Availability | Per year | Per month | Per week | Per day | Typical workload (AWS) |
|---|---|---|---|---|---|
| 99% | 3.65 d | 7.2 h | 1.68 h | 14.4 min | Batch and ETL jobs |
| 99.9% | 8.76 h | 43.2 min | 10.1 min | 1.44 min | Internal tools |
| 99.95% | 4.38 h | 21.6 min | 5.04 min | 43.2 s | Online commerce |
| 99.99% | 52.56 min | 4.32 min | 1.01 min | 8.64 s | Video delivery |
| 99.999% | 5.26 min | 25.9 s | 6.05 s | 0.86 s | ATMs, telecom |
| 99.9999% | 31.5 s | 2.59 s | 0.605 s | 0.086 s | Not in the AWS table |
What makes a number an agreement
A Service-level agreement (SLA) is where availability becomes a business promise. The vocabulary has three layers. You measure an indicator, you set an objective for it, and you sign an agreement that attaches consequences to the objective. Google's SRE book draws the line sharply: if nothing happens when the target is missed, it is an SLO, not an SLA. Azure's guidance says the same in business terms: an SLA has financial and legal implications, while SLOs are internal targets.
SLI, SLO, SLA and the consequence that binds them
- SLI (indicator)
- The measurement itself, for example the fraction of valid requests that succeeded in a 5-minute window.
- SLO (objective)
- A target value for an SLI, such as 99.95% of requests succeed each month. It is internal: missing it triggers engineering work, not payments.
- SLA (agreement)
- A contract with the customer that names one or more SLOs and the consequences of missing them. Without a consequence, it is just an SLO.
- Service credit
- The usual consequence. Amazon EC2 commits to 99.99% monthly uptime per region and refunds 10% of the bill below that, 30% below 99.0% and 100% below 95.0%. Credits are the sole remedy.
Read the Amazon EC2 SLA as a worked case. It promises 99.99% monthly uptime for instances spread across a region, but only 99.5% for a single instance. The gap is deliberate: one machine can fail, and only an architecture that spreads across failure domains earns the higher number. If AWS misses the regional promise, you receive a credit of 10%, 30% or 100% of that month's bill depending on how far uptime fell. You never receive compensation for your own lost business.
Quick check
A service promises 99.9% availability. About how much downtime does it allow per 30-day month?
Quick check
Which statement best distinguishes an SLA from an SLO?
Recall
What turns an SLO into an SLA?
Recall
Convert 99.95% to allowed downtime per 30-day month.
A request rarely touches a single service. Picture a request that first passes through Service 1, a front end promising 99.99%, and then Service 2, a database promising 99.9%. The request succeeds only if both are up at that moment. If the two fail independently, the probability that both are up is the product of their availabilities: 0.9999 × 0.999 = 0.9989001, about 99.89%.
Enters the pipeline.
Must be up.
Must also be up.
Arrives only if both succeeded.
Notice where the answer lands. It is not between the two numbers. It is below the weaker one. That result is the Compound SLA, and it holds for any chain of hard dependencies, where the failure of one stage makes the whole request fail:
The inequality follows from the arithmetic. Every is at most 1, and multiplying by a number at most 1 can only keep a value the same or lower it. AWS puts the consequence plainly: a workload can be no more available than any of its hard dependencies.
A shortcut and a warning about scale
When every unavailability is small, the product is close to one minus their sum. Here the unavailabilities are 0.01% and 0.1%, which add to 0.11%, giving roughly 99.89%. The exact value is 99.89001%, so the shortcut is good enough for an exam sanity check.
The shortcut also shows why long chains hurt. Five services at 99.9% each give 0.999⁵ ≈ 0.99501, about 99.50%, which is roughly 43.7 h of downtime a year against 8.76 h for one service. AWS gives a gentler version: a 99.99% system that hard depends on two other 99.99% systems lands at about 99.97%. Every microservice, managed database or edge gateway you add in series spends part of the budget.
Redundancy runs the arithmetic the other way
The fix for a weak link is to stop needing it to be up. Place independent replicas behind automatic failover so that any one of them can serve. Now the request fails only if all replicas are down at once, so you multiply the unavailabilities instead:
Two 99.9% replicas give 1 − 0.001² = 0.999999, or 99.9999%. AWS notes a quick way to see this when every component is all nines: count the nines, 3 + 3 = 6. The same idea explains why the Amazon EC2 SLA commits to only 99.5% for a single instance but 99.99% for instances spread across a region: the regional promise assumes your workload can survive the loss of any one instance. The two numbers are business commitments, not outputs of the formula.
| Topology | Formula | Two at 99.9% | When it applies |
|---|---|---|---|
| Series (hard dependencies) | A = ∏ Aᵢ | 99.8001% | Every service must work for the request to succeed |
| Parallel (redundant replicas) | A = 1 − ∏ (1 − Aᵢ) | 99.9999% | Any one replica can serve, and failover is automatic |
Quick check
Service 1 offers 99.99% and Service 2 offers 99.9% in series. What is the compound SLA?
Quick check
Two independent replicas, each 99.9% available, sit behind failover so either can serve. What is the effective availability?
Recall
Why can a pipeline never be more available than its weakest hard dependency?
Recall
Estimate the compound availability of 99.95% and 99.9% in series without a calculator.
A compound availability of 99.89% is still an abstract percentage. To plan operations you need it in minutes of Downtime. Its unavailability is 0.11%, or 0.0011, and the rule from the previous concepts does the rest.
Worked example
The compound 99.89% as a downtime budget
Per day
0.0011 × 86,400 s ≈ 95.0 s, which is 1 min 35 s.Per week
0.0011 × 604,800 s ≈ 665 s, which is 11 min 5 s.Per 30-day month
0.0011 × 2,592,000 s ≈ 2851 s, which is 47 min 31 s.Per 365-day year
0.0011 × 31,536,000 s ≈ 34,690 s, which is about 9 h 38 min.Result
The two services together may be down for about 48 minutes a month and nearly 10 hours a year, while still meeting the compound figure.
Online calculators such as uptime.is, the tool shown on the slide, automate exactly this multiplication. They are convenient, but they hide the period convention. uptime.is uses an average Gregorian year and month (about 365.24 d and 30.44 d), so its monthly and yearly figures come out slightly larger than the hand calculation with 30 and 365 days. The table compares the three sources.
| Period | By hand (0.11%) | uptime.is/99.89 now | Slide 67 screenshot |
|---|---|---|---|
| Day | 95.0 s = 1 min 35 s | 1m 35s | 1m 35s |
| Week | 665 s = 11 min 5 s | 11m 5.3s | 11m 5.3s |
| Month | 47 min 31 s (30 days) | 48m 13s | 47m 49s |
| Quarter | 2 h 22 min 34 s (90 days) | 2h 24m 38s | 2h 23m 27s |
| Year | 9 h 38 min (365 days) | 9h 38m 33s | 9h 33m 48s |
From allowance to error budget
Operations teams turn this allowance into a working tool called an error budget. Google's SRE book defines it as the gap between the SLO and the measured performance: the budget of how much unreliability is remaining in the period. While budget remains, the team ships releases and runs experiments, each of which risks a little downtime. When the budget is spent, releases pause and effort shifts to reliability. The Service-level agreement (SLA) sets the outer limit; the error budget is how the team manages its way to it.
Azure's reliability guidance shows the same thinking from the design side. In its worked example the composite of a workload's components comes to about 99.45%, roughly 4 h of downtime a month. The business then signs a 99.90% SLA anyway, consciously accepting the risk of paying credits. That gap between what the architecture supports and what the contract promises is a business decision, and it is exactly what the compound arithmetic lets you see before you sign.
Recall
Three services at 99.95%, 99.9% and 99.99% run in series. Roughly how much downtime does the chain allow per 30-day month?
Recall
Why does uptime.is report 48m 13s per month at 99.89% when a hand calculation gives 47m 31s?
Recap
If you remember nothing else
- Availability = uptime / (uptime + downtime). Allowed downtime = (1 − A) × T for any period T.
- Each extra nine cuts the downtime budget tenfold: 99.9% allows 8.76 h per year and 43.2 min per 30-day month.
- An SLA is a contract with consequences (Amazon EC2 pays service credits). An SLO is an internal target, and an SLI is the measurement behind both.
- Hard dependencies in series multiply: 99.99% × 99.9% ≈ 99.89%, which is below even the weakest link.
- Independent redundant replicas combine as 1 − ∏(1 − Aᵢ): two 99.9% replicas give 99.9999%.
- Downtime calculators differ by period convention, so check them. The calculator screenshot on slide 67 is not even consistent with itself.
Sources
- Reliability Pillar: AvailabilityDocsAmazon Web Services, Well-Architected FrameworkDefinition, time and request based formulas, MTBF and MTTR estimate, hard dependencies, redundancy, tier examples.(opens in a new tab)
- Availability with dependenciesDocsAmazon Web Services, Availability and Beyond whitepaperProduct formula, rough-bound caveat, rules on reducing and choosing dependencies.(opens in a new tab)
- Amazon Compute Service Level AgreementDocsAmazon Web Services99.99% regional and 99.5% single-instance commitments, service credit tiers.(opens in a new tab)
- Architecture strategies for defining reliability targetsDocsMicrosoft Learn, Azure Well-Architected FrameworkSLA versus SLO, composite SLO as a product, the 99.45% composite and 99.90% SLA example.(opens in a new tab)
- Site Reliability Engineering, Chapter 3: Embracing RiskBookGoogle, O'Reilly 2016Time-based and request-based availability, nines as orders of magnitude, error budgets.(opens in a new tab)
- Site Reliability Engineering, Chapter 4: Service Level ObjectivesBookGoogle, O'Reilly 2016Definitions of SLI, SLO and SLA.(opens in a new tab)
- Site Reliability Engineering, Appendix A: Availability TableBookGoogle, O'Reilly 2016(opens in a new tab)
- NIST SP 500-307: Cloud Computing Service Metrics DescriptionDocsNIST, 2018Background on measurable cloud metrics used in SLAs.(opens in a new tab)
- SLA and uptime calculator: 99.89%Articleuptime.isThe calculator shown on slide 67.(opens in a new tab)