Majid Al-RaimiService-level agreements and availability

COE 558Lecture 02Part 10

Service-level agreements and availability

Defines availability, maps SLA nines to allowed downtime, and shows how serial services combine into a compound SLA.

Concepts
4
Slides
63-68
Reading
24 min
Understood
0/4 concepts

Why this part matters

The earlier parts of this lecture asked where a service should run and who should own the hardware. This last part asks a sharper question: how often will that service actually be there when a user calls it? Every Cloud service provider (CSP) sells the answer as a single number with a row of nines, and every architecture you design chains several such numbers together.

Turning nines into minutes of downtime, and composing them correctly across a pipeline, is a calculation you should expect on the exam. You will also need it the moment your research system spans an edge site, a fog node and a cloud region: each hop is another dependency, and each dependency spends part of your reliability budget.

By the end you can

  1. Compute availability from uptime and downtime, from successful and valid requests, or from MTBF and MTTR.
  2. Convert any number of nines into allowed downtime per day, week, month and year, and state the period convention you used.
  3. Compute the compound SLA of services in series and contrast it with the availability of redundant replicas in parallel.
  4. Distinguish SLI, SLO and SLA, and read the commitments and service credits of a real provider SLA.

Take a campus course-registration API and watch it for a 30-day month, which is 720 h. During that month it is unreachable twice, for a total of 2 h. For the other 718 h it answered requests. The fraction of the month it could answer is 718 / 720 = 0.9972, or 99.72%. That fraction is the service's Availability.

The general rule only names the two pieces. Uptime is the time the application can serve requests, and Downtime is the time it cannot. Availability is the share of uptime in the total:

A=UptimeUptime+DowntimeA=\frac{\text{Uptime}}{\text{Uptime}+\text{Downtime}}
Availability as the share of total time the service can answer

The denominator is simply the length of the observation period, TT. Rearranging gives the form you will use most often, because it turns a promised availability into an allowance of failure: D=(1−A)×TD=(1-A)\times T. For the registration API, (1 − 0.9972) × 720 h ≈ 2 h, which is exactly where we started.

Counting requests instead of hours

Large providers rarely have a clean on or off switch. A service may be up for most users while one shard fails. So they often measure availability by requests rather than by time: successful requests divided by valid requests. Google's SRE book gives a concrete case. A system that serves 2.5 million requests a day with a daily target of 99.99% may fail up to 250 of them. AWS combines both views: it measures success in windows of 1 or 5 minutes and averages the windows into a monthly uptime percentage.

Estimating availability before you have measurements

At design time there is no uptime log yet. Instead you can estimate availability from two failure statistics: the mean time between failures (MTBF) and the mean time to recover (MTTR).

Aest=MTBFMTBF+MTTRA_{est}=\frac{\text{MTBF}}{\text{MTBF}+\text{MTTR}}
Availability estimated from failure frequency and recovery time

Worked example

A component that fails every 150 days

  1. Put both statistics in one unit

    MTBF is 150 d = 3600 h. MTTR is 1 h.
  2. Apply the estimate

    A = 3600 / (3600 + 1) = 3600 / 3601 ≈ 0.99972.
  3. Result

    About 99.97%. The formula shows the two levers you have: fail less often (raise MTBF) or recover faster (cut MTTR). Halving the recovery time to 30 min buys as much as doubling the time between failures.

Recall

A component has an MTBF of 30 days and an MTTR of 2 hours. What is its availability, and which two levers raise it?

720 / 722 ≈ 99.72%. Raise MTBF (fail less often) or cut MTTR (recover faster, for example with automatic failover).

From nines to a downtime budget

Start with the most common promise in cloud marketing, 99.9%, often called “three nines”. Its unavailability is 1 − 0.999 = 0.001. Multiply that by a period and you get the Downtime it allows: 0.001 × 8760 h = 8.76 h a year, 0.001 × 720 h = 43.2 min in a 30-day month, and 0.001 × 168 h ≈ 10.1 min a week.

Now add one more nine. An Availability of 99.99% has an unavailability of 0.0001, ten times smaller, so every downtime figure shrinks tenfold: 52.56 min a year, 4.32 min a month. That is the whole rule. Google's SRE book calls each additional nine “an order of magnitude improvement”, and the cost of getting there typically rises much faster than tenfold.

D=(1−A)×TD=(1-A)\times T
Downtime allowed by availability A over a period T
One year per row. Each extra nine cuts the real downtime tenfold (labels on the right); the slivers are not to scale and shrink only about 2.5 times per row so the smallest stay visible.

The table below repeats the calculation for each common tier. Like the lecture slide, it assumes a 365-day year, a 30-day month and a 7-day week. Every cell on the slide recomputes correctly with those conventions, and the figures agree with the availability table in the Google SRE book. The last column gives the example workloads AWS associates with each tier, so you can feel what a nine buys.

AvailabilityPer yearPer monthPer weekPer dayTypical workload (AWS)
99%3.65 d7.2 h1.68 h14.4 minBatch and ETL jobs
99.9%8.76 h43.2 min10.1 min1.44 minInternal tools
99.95%4.38 h21.6 min5.04 min43.2 sOnline commerce
99.99%52.56 min4.32 min1.01 min8.64 sVideo delivery
99.999%5.26 min25.9 s6.05 s0.86 sATMs, telecom
99.9999%31.5 s2.59 s0.605 s0.086 sNot in the AWS table
Allowed downtime by availability tier

What makes a number an agreement

A Service-level agreement (SLA) is where availability becomes a business promise. The vocabulary has three layers. You measure an indicator, you set an objective for it, and you sign an agreement that attaches consequences to the objective. Google's SRE book draws the line sharply: if nothing happens when the target is missed, it is an SLO, not an SLA. Azure's guidance says the same in business terms: an SLA has financial and legal implications, while SLOs are internal targets.

SLI, SLO, SLA and the consequence that binds them

SLI (indicator)
The measurement itself, for example the fraction of valid requests that succeeded in a 5-minute window.
SLO (objective)
A target value for an SLI, such as 99.95% of requests succeed each month. It is internal: missing it triggers engineering work, not payments.
SLA (agreement)
A contract with the customer that names one or more SLOs and the consequences of missing them. Without a consequence, it is just an SLO.
Service credit
The usual consequence. Amazon EC2 commits to 99.99% monthly uptime per region and refunds 10% of the bill below that, 30% below 99.0% and 100% below 95.0%. Credits are the sole remedy.

Read the Amazon EC2 SLA as a worked case. It promises 99.99% monthly uptime for instances spread across a region, but only 99.5% for a single instance. The gap is deliberate: one machine can fail, and only an architecture that spreads across failure domains earns the higher number. If AWS misses the regional promise, you receive a credit of 10%, 30% or 100% of that month's bill depending on how far uptime fell. You never receive compensation for your own lost business.

Quick check

A service promises 99.9% availability. About how much downtime does it allow per 30-day month?

Quick check

Which statement best distinguishes an SLA from an SLO?

Recall

What turns an SLO into an SLA?

A contract with consequences, such as service credits. Amazon EC2 refunds 10%, 30% or 100% of the monthly bill depending on how far uptime fell below its 99.99% regional commitment.

Recall

Convert 99.95% to allowed downtime per 30-day month.

Unavailability is 0.0005. A 30-day month is 43,200 min, so 0.0005 × 43,200 = 21.6 min.

A request rarely touches a single service. Picture a request that first passes through Service 1, a front end promising 99.99%, and then Service 2, a database promising 99.9%. The request succeeds only if both are up at that moment. If the two fail independently, the probability that both are up is the product of their availabilities: 0.9999 × 0.999 = 0.9989001, about 99.89%.

Request
client

Enters the pipeline.

Service 1
99.99%

Must be up.

Service 2
99.9%

Must also be up.

Response
compound ≈ 99.89%

Arrives only if both succeeded.

A request that needs both services: the pipeline is only up when every stage is up

Notice where the answer lands. It is not between the two numbers. It is below the weaker one. That result is the Compound SLA, and it holds for any chain of hard dependencies, where the failure of one stage makes the whole request fail:

Aseries=∏i=1nAi  ≤  min⁡iAiA_{\text{series}}=\prod_{i=1}^{n} A_i \;\le\; \min_i A_i
Hard dependencies in series multiply, so the chain is never better than its weakest stage

The inequality follows from the arithmetic. Every AiA_i is at most 1, and multiplying by a number at most 1 can only keep a value the same or lower it. AWS puts the consequence plainly: a workload can be no more available than any of its hard dependencies.

Two services in series. The compound bar (teal) settles just short of the weaker service, on a zoomed 99.8% to 100% scale.

A shortcut and a warning about scale

When every unavailability is small, the product is close to one minus their sum. Here the unavailabilities are 0.01% and 0.1%, which add to 0.11%, giving roughly 99.89%. The exact value is 99.89001%, so the shortcut is good enough for an exam sanity check.

The shortcut also shows why long chains hurt. Five services at 99.9% each give 0.999⁵ ≈ 0.99501, about 99.50%, which is roughly 43.7 h of downtime a year against 8.76 h for one service. AWS gives a gentler version: a 99.99% system that hard depends on two other 99.99% systems lands at about 99.97%. Every microservice, managed database or edge gateway you add in series spends part of the budget.

Redundancy runs the arithmetic the other way

The fix for a weak link is to stop needing it to be up. Place independent replicas behind automatic failover so that any one of them can serve. Now the request fails only if all replicas are down at once, so you multiply the unavailabilities instead:

Aparallel=1−∏i=1n(1−Ai)A_{\text{parallel}}=1-\prod_{i=1}^{n}(1-A_i)
Independent redundant replicas: the system fails only when every replica fails

Two 99.9% replicas give 1 − 0.001² = 0.999999, or 99.9999%. AWS notes a quick way to see this when every component is all nines: count the nines, 3 + 3 = 6. The same idea explains why the Amazon EC2 SLA commits to only 99.5% for a single instance but 99.99% for instances spread across a region: the regional promise assumes your workload can survive the loss of any one instance. The two numbers are business commitments, not outputs of the formula.

TopologyFormulaTwo at 99.9%When it applies
Series (hard dependencies)A = ∏ Aᵢ99.8001%Every service must work for the request to succeed
Parallel (redundant replicas)A = 1 − ∏ (1 − Aᵢ)99.9999%Any one replica can serve, and failover is automatic
Series versus parallel, for two components at 99.9%
Top: in series, one failed service cuts the path. Bottom: in parallel, the failed replica dims and traffic reroutes through the other.

Quick check

Service 1 offers 99.99% and Service 2 offers 99.9% in series. What is the compound SLA?

Quick check

Two independent replicas, each 99.9% available, sit behind failover so either can serve. What is the effective availability?

Recall

Why can a pipeline never be more available than its weakest hard dependency?

Its availability is a product of numbers no greater than 1. Each factor can only keep the value the same or lower it, so the product is at most min⁡iAi\min_i A_i.

Recall

Estimate the compound availability of 99.95% and 99.9% in series without a calculator.

Add the unavailabilities: 0.05% + 0.1% = 0.15%, so about 99.85%. The exact value is 0.9995 × 0.999 = 0.9985005, or 99.85005%.

A compound availability of 99.89% is still an abstract percentage. To plan operations you need it in minutes of Downtime. Its unavailability is 0.11%, or 0.0011, and the rule from the previous concepts does the rest.

Worked example

The compound 99.89% as a downtime budget

  1. Per day

    0.0011 × 86,400 s ≈ 95.0 s, which is 1 min 35 s.
  2. Per week

    0.0011 × 604,800 s ≈ 665 s, which is 11 min 5 s.
  3. Per 30-day month

    0.0011 × 2,592,000 s ≈ 2851 s, which is 47 min 31 s.
  4. Per 365-day year

    0.0011 × 31,536,000 s ≈ 34,690 s, which is about 9 h 38 min.
  5. Result

    The two services together may be down for about 48 minutes a month and nearly 10 hours a year, while still meeting the compound figure.

Online calculators such as uptime.is, the tool shown on the slide, automate exactly this multiplication. They are convenient, but they hide the period convention. uptime.is uses an average Gregorian year and month (about 365.24 d and 30.44 d), so its monthly and yearly figures come out slightly larger than the hand calculation with 30 and 365 days. The table compares the three sources.

PeriodBy hand (0.11%)uptime.is/99.89 nowSlide 67 screenshot
Day95.0 s = 1 min 35 s1m 35s1m 35s
Week665 s = 11 min 5 s11m 5.3s11m 5.3s
Month47 min 31 s (30 days)48m 13s47m 49s
Quarter2 h 22 min 34 s (90 days)2h 24m 38s2h 23m 27s
Year9 h 38 min (365 days)9h 38m 33s9h 33m 48s
Downtime allowed at 99.89%: by hand, on uptime.is today, and on the slide's screenshot

From allowance to error budget

Operations teams turn this allowance into a working tool called an error budget. Google's SRE book defines it as the gap between the SLO and the measured performance: the budget of how much unreliability is remaining in the period. While budget remains, the team ships releases and runs experiments, each of which risks a little downtime. When the budget is spent, releases pause and effort shifts to reliability. The Service-level agreement (SLA) sets the outer limit; the error budget is how the team manages its way to it.

Azure's reliability guidance shows the same thinking from the design side. In its worked example the composite of a workload's components comes to about 99.45%, roughly 4 h of downtime a month. The business then signs a 99.90% SLA anyway, consciously accepting the risk of paying credits. That gap between what the architecture supports and what the contract promises is a business decision, and it is exactly what the compound arithmetic lets you see before you sign.

Recall

Three services at 99.95%, 99.9% and 99.99% run in series. Roughly how much downtime does the chain allow per 30-day month?

Unavailabilities add to about 0.0016 (exact 0.1599%), so 0.0016 × 43,200 min ≈ 69 min.

Recall

Why does uptime.is report 48m 13s per month at 99.89% when a hand calculation gives 47m 31s?

It uses an average Gregorian month of about 30.44 days instead of 30 days. The formula (1−A)×T(1-A)\times T is the same; only T differs.

Recap

If you remember nothing else

  • Availability = uptime / (uptime + downtime). Allowed downtime = (1 − A) × T for any period T.
  • Each extra nine cuts the downtime budget tenfold: 99.9% allows 8.76 h per year and 43.2 min per 30-day month.
  • An SLA is a contract with consequences (Amazon EC2 pays service credits). An SLO is an internal target, and an SLI is the measurement behind both.
  • Hard dependencies in series multiply: 99.99% × 99.9% ≈ 99.89%, which is below even the weakest link.
  • Independent redundant replicas combine as 1 − ∏(1 − Aᵢ): two 99.9% replicas give 99.9999%.
  • Downtime calculators differ by period convention, so check them. The calculator screenshot on slide 67 is not even consistent with itself.

Sources