COE 558Lecture 02Full guide
E2C continuum and cloud computing
The whole lecture on one page, taught concept by concept. Work through the parts in order, mark each concept once you understand it, and open the slide chips when you want the original slides.
- Parts
- 10
- Concepts
- 50
- Slides
- 68
- Reading
- 300 min
Part 01: The latency continuum and the birth of the fog
Why the cloud cannot simply move closer to the client, how mini data centers along the latency continuum form the fog, and which points need networking at all.
5 concepts, slides 1-9
Why this part matters
Every system you design in COE 558, and in your research, starts with a placement decision: where does each piece of work run? Lecture 01 showed that latency has a floor set by physics. This part shows how the industry answered that floor without giving up the economics of the cloud. It placed small copies of the data center, the fog, along the path to the user, and it uses a simple scope rule to decide which of those points need networking at all. The classify-and-justify question (edge, fog or cloud?) is the most likely exam item from this section, so every concept below ends on that skill.
By the end you can
- Explain the E2C continuum as a continuum of latencies, and estimate the round-trip floor for a placement from its distance.
- Explain why the cloud stays centralized (economies of scale) and how mini data centers and CDNs such as Netflix Open Connect bring service closer anyway.
- Define the fog, its mini data centers and their compute, storage and networking services, and tell fog apart from edge.
- Use a function's scope to decide whether it needs networking, and place tasks on edge (LAN), fog (WAN) or cloud (Internet) with a justification.
Picture an autonomous car with three very different jobs. First, it must detect an obstacle and brake, and that answer has to arrive within milliseconds. Second, it wants a good route, which means combining traffic reports from many cars in the same region. Third, its driving model must improve over time, and that means training on months of history from the whole fleet. No single computer suits all three. The obstacle check belongs right next to the sensors. Route optimization belongs somewhere that sees the whole region. Training belongs where there is enormous storage and compute. The three places sit at small, medium and large distance from the car.
The lecture turns that intuition into a definition. The Edge-to-cloud (E2C) continuum is a continuum of latencies between a client and a server. The Edge layer, the Fog layer and the Cloud layer are not separate worlds. They are regions of one range, ordered by how long a request takes to reach them and come back. Moving away from the client adds latency, and in exchange it adds capacity: more processors, more storage, a wider view of the data.
| Task | Deadline | Data it needs | Layer |
|---|---|---|---|
| Obstacle detection | A few ms | This car's own sensors | Edge |
| Route optimization | Seconds | Traffic from many cars in a region | Fog |
| Model training | Hours to days | Months of history from the whole fleet | Cloud |
The floor that distance sets
Recall
From Lecture 01: what floor does distance set on RTT, and why?
Glass has a refractive index of about 1.4 to 1.6 (typically around 1.5), so light in fiber moves at roughly 200,000 km/s. Applied to the continuum, that floor reads:
Worked example
Physics floor for three placements
Cloud region 4,000 km away
One way takes 4,000 km ÷ 200,000 km/s = 0.02 s = 20 ms, so the round trip is at least 40 ms before any queuing, routing or processing.Fog site 50 km away
One way takes 0.25 ms, so the round trip is at least 0.5 ms.Edge on the same LAN, about 100 m away
One way takes 0.5 µs, so the round trip is at least 1 µs. Distance has effectively vanished; only processing remains.Compare with a real measurement
Satyanarayanan reports an average round trip of 74 ms from 260 vantage points to their best Amazon EC2 region, and a wireless first hop adds more on top. Real paths sit well above the floor because every hop adds queuing and routing delay.Result
Moving a server from 4,000 km to 50 km lowers the floor 80-fold. That is the whole argument for the continuum in one number.
To try your own distances and media, use the latency calculator in Lecture 01 part 07, which also explains why real round trips cost several times this bound.
Recall
What exactly is "continuous" in the E2C continuum?
Recall
What is the minimum RTT to a fog site 300 km away over fiber?
A multiplayer game hosted in a distant cloud region lags: every shot travels thousands of kilometres and back. The obvious fix is to move the cloud next to the players. The lecture's answer is a flat no, and the reason is worth understanding, because it shapes everything that follows.
A cloud data center is a very complex system, it is very expensive, and it is centralized on purpose. Building one is, in the words of the Berkeley "Above the Clouds" report, a hundred-million-dollar undertaking. The payoff of that size is Economies of scale: a very large data center buys network bandwidth, storage and administration at a fraction of the unit cost a medium one pays. Satyanarayanan puts it the same way: centralization exploits economies of scale to lower the marginal cost of system administration and operations. Split one giant site into a thousand neighborhood sites and you throw that advantage away.
| Resource | Medium datacenter | Very large datacenter | Ratio |
|---|---|---|---|
| Network | $95 per Mbit/s per month | $13 | 7.1 |
| Storage | $2.20 per GB per month | $0.40 | 5.7 |
| Administration | about 140 servers per admin | more than 1,000 | 7.1 |
So the cloud stays where it is, and something smaller goes out instead. The lecture calls it a mini cloud, or Mini data center, and the key move is that you can place several of them at different points on the latency continuum. Each one is close enough to cut latency for nearby clients, while the big data center keeps its economics.
Player, viewer or device
Cached content and services
Serves a whole metro area
Authoritative data and heavy jobs
The pattern in production: Netflix Open Connect
A Content delivery network (CDN) is the clearest working example. Netflix started its Open Connect program in 2011. It places its own caching appliances (Open Connect Appliances, OCAs) at internet exchange points and, free of charge, inside the networks of qualifying ISPs; it now works with more than a thousand ISP partners. The ISP provides power, space and connectivity. Because Netflix can predict what its members will watch, it pre-fills each appliance during off-peak "fill windows", so at peak time the popular titles are already a few hops from the viewer.
Recall
Give two reasons the cloud cannot simply be moved next to clients, and the alternative.
Quick check
Why does the lecture reject moving the whole cloud closer to clients?
In a smart city, small computers sit in cabinets at road junctions. Each one reads the cameras and sensors at its junction, notices a queue building or a pedestrian waiting, and changes the signal timing in real time, with no round trip to a distant cloud. Bonomi and colleagues at Cisco used almost the same example when they introduced the idea in 2012: a smart traffic light node that reads local sensors detecting pedestrians and bikers.
The layer of such mini data centers gets a name: the fog. A Mini data center is a scaled-down version of a big data center that bridges the gap between the edge and the cloud, and it offers all three classic services: compute, storage and networking. Bonomi defined fog computing as "a highly virtualized platform that provides compute, storage, and networking services between end devices and traditional Cloud Computing Data Centers." The name is a joke with a point: the fog is a cloud close to the ground.
Bonomi listed what makes the fog distinct from the cloud above it:
- Low latency and location awareness: a node knows where it is and serves who is near.
- Wide-spread geographic distribution and a very large number of nodes.
- Support for mobility, wireless access and real-time streaming.
- Heterogeneity: nodes come in many shapes and from many vendors.
The fog is also hierarchical. Nodes form tiers, and in Bonomi's words, "the higher the tier, the wider the geographical coverage, and the longer the time scale." The lowest tier runs control loops in milliseconds to sub-seconds; higher tiers work over seconds to minutes and even days, and the cloud keeps data for months and years. NIST SP 500-325 later made the building blocks concrete: fog nodes can be physical (gateways, switches, routers, servers) or virtual (virtualized switches, virtual machines, cloudlets). It also names mist computing, a lightweight and optional form of fog at the very edge. In 2018 IEEE adopted the OpenFog Reference Architecture as standard IEEE 1934-2018, a sign that the idea had moved from paper to practice.
| Aspect | Fog | Edge |
|---|---|---|
| Structure | Hierarchical, multi-layer architecture | Limited to a small number of peripheral devices |
| Services | Computation and networking, plus storage, control and data-processing acceleration | Specific applications in a fixed logic location, with a direct transmission service |
| Location | Between smart end-devices and centralized cloud services | The end-device layer itself |
Recall
Give two differences between fog and edge according to NIST.
Consider a fleet of autonomous delivery drones. When one drone's battery runs low it must return to base. That decision uses nothing but the drone's own state, and it must work even if every radio link is down, so it runs on the drone's own compute with no networking at all. When several drones fly in the same area, they must agree on flight paths to avoid collisions. Now the data spans several parties in one region, so the drones talk to a fog ground station, and networking is essential. Finally, the cloud combines data from every drone in the fleet to improve routes over time. That function's scope is global.
The lecture asks a sharp question: do we really need the networking service at every point on the latency continuum? The drone example answers it. Whether a function needs networking depends on its scope, meaning how many parties and how much area its data spans, not on whether the layer is called edge, fog or cloud. Walking from the cloud toward the device (the numbered points 1, 2 and 3 on the slides), more functions can run locally and depend less on remote communication, because the functions that live there have narrower scope.
| Task | Scope | Layer | Needs networking? |
|---|---|---|---|
| Low battery, return to base | One drone | Edge, on the drone itself (no network hop) | No |
| Coordinate flight paths | Several drones in one area | Fog ground station | Yes |
| Improve routes from fleet history | Every drone, everywhere | Cloud | Yes |
Here the device itself plays the edge role, the NIST usage flagged in concept 1.
Worked example
Classify a function by scope
Whose data does it need?
One device, a neighborhood of devices, or everyone. This sets the scope and therefore whether networking is needed.What is its deadline?
Milliseconds push it toward the client; minutes or days let it go far.Place it at the nearest point that covers that scope
A one-device, tight-deadline function lands on the edge. A regional, multi-party function lands in the fog. A global, slow function lands in the cloud.Result
Scope decides networking and coverage; deadline decides how near. The answer is the closest layer that satisfies both.
Research supports the same rule. Satyanarayanan shows that analyzing raw video on a nearby cloudlet means only the "(much smaller) extracted information and metadata must be transmitted to the cloud", and the cloudlet can enforce privacy policies before anything leaves. Bonomi's lowest fog tier runs machine-to-machine control loops in milliseconds and filters data before passing it up. In both cases, local functions stay local, and only results with wider scope travel.
Recall
What decides whether a function needs networking?
Quick check
A smart thermostat must cut the heater when the room overheats, even offline. Where should it run?
Take the drone tasks one last time and map them onto networks. Local work runs on the edge, on the device or inside its local area network. Collision coordination runs in the fog, across a wide area network run by an operator. Fleet-wide learning runs in the cloud, reached over the public Internet. That mapping is the lecture's summary picture, and it is the one to memorize.
The edge lives on the LAN and offers compute and storage. The fog lives on the WAN and offers compute, storage and networking. The cloud lives on the Internet and is the full data center, offering every service at global scale. Each has a priority: edge prioritizes autonomy, fog prioritizes coordination, and cloud prioritizes global connectivity. Read together with the scope rule, this is the whole Edge-to-cloud (E2C) continuum in one table.
| Layer | Network span | Services offered | Priority | Typical RTT floor | Example task |
|---|---|---|---|---|---|
| Edge | LAN | Compute, storage | Autonomy | ~1 µs at 100 m | Obstacle detection, local video analytics |
| Fog | WAN | Compute, storage, networking | Coordination | ~0.5 ms at 50 km | Collision coordination, traffic signals |
| Cloud | Internet | Full data center (all services) | Global connectivity | ~40 ms at 4,000 km | Fleet-wide model training |
Where does an edge layer meet the fog? Through an Edge gateway, a device that does local processing for the edge devices behind it and connects them upward to fog nodes. The edge box omits networking because the functions placed there have device-local scope, as slides 7 and 8 argued; the gateway's uplink gives the edge connectivity, but the edge offers no networking service to others.
Recall
Fill in slide 9: network span, services and priority for edge, fog and cloud.
Quick check
On slide 9, which layer offers compute and storage but no networking service?
Quick check
Traffic lights at twelve nearby junctions must share queue lengths to time a green wave. Which placement fits best?
Recap
If you remember nothing else
- The E2C continuum is a range of latencies: edge small, fog medium, cloud large, with capacity rising in the same direction.
- Distance sets a hard floor: RTT ≥ 2d/v, with v ≈ 200,000 km/s in fiber.
- The cloud stays centralized because scale cuts unit costs by roughly 5 to 7 times. Mini data centers bring service closer instead.
- The fog is the layer of mini data centers. It offers compute, storage and networking, is geo-distributed and hierarchical, and is "a cloud close to the ground".
- Fog is not edge: fog is hierarchical and multi-layer; edge is limited to a few peripheral devices running fixed applications.
- Networking needs follow the function's scope, not the layer's name.
- Edge (LAN) offers compute and storage and prioritizes autonomy; fog (WAN) adds networking and prioritizes coordination; cloud (Internet) prioritizes global connectivity.
Sources
- NIST SP 500-325: Fog Computing Conceptual Model (Iorga et al., March 2018)DocsNISTFog definition, fog nodes, mist, Annex A fog against edge(opens in a new tab)
- Fog Computing and Its Role in the Internet of Things (Bonomi, Milito, Zhu, Addepalli, MCC 2012)PaperACMFog definition, characteristics, tiers and time scales, traffic light example(opens in a new tab)
- The Emergence of Edge Computing (Satyanarayanan, IEEE Computer 50(1), 2017)PaperIEEE Computer Society74 ms RTT to EC2, economies of scale, edge analytics, masking outages(opens in a new tab)
- The Case for VM-Based Cloudlets in Mobile Computing (Satyanarayanan, Bahl, Cáceres, Davies, IEEE Pervasive Computing 2009)PaperIEEECloudlet as a data center in a box with cached state(opens in a new tab)
- Above the Clouds: A Berkeley View of Cloud Computing (Armbrust et al., UCB/EECS-2009-28)PaperUC BerkeleyTable 2, economies of scale in 2006(opens in a new tab)
- Netflix Open Connect OverviewDocsNetflixOCAs at IXPs and in ISPs, fill windows(opens in a new tab)
- Netflix Open ConnectDocsNetflixMore than a thousand ISP partners(opens in a new tab)
- High Performance Browser Networking: Primer on Latency and Bandwidth (Grigorik)BookO'ReillySpeed of light in fiber, refractive index(opens in a new tab)
- IEEE adopts OpenFog Consortium's reference architecture (June 2018)ArticleRCR Wireless NewsIEEE 1934-2018(opens in a new tab)
Part 02: Roles of the edge, fog and cloud layers
What each layer of the continuum provides, why the fog helps, and two worked architectures (IoT hierarchy and a vehicle platoon).
5 concepts, slides 10-14
Why this part matters
Part 01 named the layers. This part says what each layer is for.
In the exam, in the system designs of your research, and in real IoT or vehicular deployments, you justify where a task runs by the role of the layer: the edge serves the client, the fog aggregates many edges, the cloud gives scale. Two worked architectures, an IoT tree and a vehicle platoon, close the part. They are the templates you will reuse whenever someone asks you to place a system on the continuum.
By the end you can
- Explain the three benefits of a fog layer that serves many edges, and quantify the bandwidth saving in a concrete scenario.
- Describe the four roles of the edge layer, including why it hides the fog from clients.
- List the cloud's big compute, big storage and big networking, with a real service for each.
- Map the devices of an IoT hierarchy and of a vehicle platoon to edge, fog and cloud, and justify each placement with latency.
- Recognise how NIST SP 500-325 and this course draw the edge and fog boundary differently.
A city installs 200 traffic cameras, and each one streams video at 4 Mbps. If every stream goes straight to the cloud, 800 Mbps of raw video crosses the wide-area network every second, around the clock, mostly showing empty roads.
Now put one small server room in the district, between the cameras and the cloud. It runs vehicle detection on all 200 streams and forwards only the events: one 2 kB record per camera per second saying what passed and how fast.
Worked example
Camera bandwidth at the fog
Raw traffic without a fog
200 × 4 Mbps = 800 Mbps into the cloud.Event traffic after the fog
Each camera now sends 2 kB × 8 = 16 kbit per second, so the whole district sends 200 × 16 kbit/s = 3.2 Mbps.Result
About 800 / 3.2 = 250× less traffic leaves the district, and the cloud still learns everything it needs for city-wide planning.
The numbers are not a toy. Satyanarayanan estimates that only 12,000 users streaming 1080p video would already need a 100 Gbps link into the cloud, and a million users would need 8.5 Tbps. His cloudlets exist largely to “reduce ingress bandwidth into the cloud”.
The rule: a shared layer between edges and cloud
The server room is what this course calls the Fog layer: a layer that provides compute, storage and networking at critical points closer to the client on the Edge-to-cloud (E2C) continuum. Physically it is usually a Mini data center, the "cloud close to the ground" that part 01 introduced with Bonomi's definition. Part 01 asked what the fog is; this concept asks what it is for.
The defining fact is that one Fog node serves many edges. Everything useful about the fog follows from that sharing:
- Less edge-to-cloud transfer. The fog filters, aggregates and answers locally, so less data crosses the long, expensive path. That lowers Latency for the requests it answers and frees bandwidth for everything else, as the camera example showed.
- Data geo-restriction. Because the fog sits inside a region, personal data can be processed there and never leave. This is Data geo-restriction, and it exists for legal and privacy compliance.
- Shorter edge-to-edge paths. Two edges under the same fog can talk through it. Their traffic turns around in the district instead of travelling to a distant data center and back.
Why geo-restriction is a legal requirement, not a preference
Data protection laws restrict where personal data may go. The EU GDPR, Chapter V, Article 44, allows a transfer to a third country only under the conditions that chapter lays down. Saudi Arabia's Personal Data Protection Law, Article 29, governs transfer outside the Kingdom, and SDAIA publishes standard contractual clauses for such transfers. A Riyadh hospital can therefore keep raw scans on a fog node inside the Kingdom and send only anonymised aggregates onward to a cloud region elsewhere.
Satyanarayanan adds two related benefits. A cloudlet can enforce its owner's privacy policy before any data is released to the cloud, and it can mask a transient cloud outage by keeping local services running until the link returns.
Recall
Name three benefits the fog gets from serving many edges.
Quick check
A Riyadh hospital's raw scans may not leave Saudi Arabia, yet it wants cloud analytics. Which fog benefit makes this possible?
A robot arm on a factory floor must recognise the part in front of it. Its own processor is too weak for the detection model, so it sends each frame to a gateway box on the factory LAN and gets back a label. The robot knows one address. It never learns whether the gateway answered by itself or passed the hard frames to a fog server two buildings away.
The rule: four roles at the boundary
That gateway box is the Edge layer. In this course the edge has four roles:
- WAN boundary. It sits on the physical boundary of the wide-area network that extends the cloud, the last stop before the client's own network.
- Offload target. The client can hand it compute jobs it cannot or should not run itself, such as the robot's detection model.
- Fast small-data processing. It handles small amounts of data at reasonable speed, close enough that the round trip stays short.
- Gateway that hides the fog. It is the client's single door into the rest of the continuum, and the fog behind it stays invisible.
The last role has a precise parallel in HTTP. RFC 9110, §3.7, defines a gateway, also called a reverse proxy, as an intermediary that “acts as an origin server for the outbound connection but translates received requests and forwards them inbound to another server or servers”. The client believes it is talking to the server. The edge plays exactly that part for the fog.
Hiding matters because the fog changes. Operators add nodes, fail over to a spare, or move a service to a less loaded server. If clients addressed fog nodes directly, every such change would have to be pushed to every robot, sensor and phone. With the edge in front, only the edge's forwarding changes and the clients never notice.
Real products follow this shape. AWS IoT Greengrass runs on a local gateway and lets devices “act locally on the data that they generate”, “filter and aggregate device data” and “react autonomously to local events”, while staying connected to AWS for the rest.
| Question | This course | NIST SP 500-325 |
|---|---|---|
| What “edge” names | The layer on the physical boundary of the WAN | The layer of end devices and their users |
| Where a gateway sits | In the edge layer (slide 13) | Among the fog nodes |
| What the client talks to | The edge, which hides the fog | A nearby fog node, often a gateway |
| When to use it | Exam answers for this course | Papers and standards that cite NIST |
Recall
What does “the edge hides the fog from the client” mean, and why is it useful?
A city wants to predict tomorrow's traffic from two years of data from every intersection. Training that model needs thousands of GPUs for days and petabytes of history in one place. No district fog room or edge gateway has either, and no single district has all the data.
The rule: scale in three forms
That job belongs to the Cloud layer, which provides compute, storage and networking at a much bigger scale:
- Big compute means high-performance computing. AWS advertises that you can “scale HPC applications to thousands of CPUs and GPUs with Elastic Fabric Adapter”.
- Big storage means big data that lives for a long time. Bonomi describes the cloud as the “repository for data that has a permanence of months and years”, and Amazon S3 is a real service built for that kind of long-lived data at petabyte scale.
- Big networking means connectivity between clouds. Google Cross-Cloud Interconnect gives dedicated links from Google Cloud to AWS, Azure, OCI and Alibaba Cloud at 10 Gbps, 100 Gbps or 400 Gbps (the 400 Gbps option only to AWS and OCI).
Putting the three layers side by side
With all three roles in hand, the layers line up along two axes at once. Going up from edge to fog to cloud, the area served grows, and so does how long the data is kept and acted on. Bonomi summarises the interplay as “Fog localization, and Cloud globalization”.
| Layer | Primary role | Scope served | Data horizon | Example |
|---|---|---|---|---|
| Edge | Immediate processing near the client | One client or one site | Milliseconds (act and discard) | Wake-word detection on a smart speaker |
| Fog | Intermediate layer for a local region | Many edges in a district | Seconds to days | Traffic management over nearby vehicle data |
| Cloud | Centralized, large-scale processing | Everyone, globally | Months to years (long-term retention) | Training an AI model in a data center |
The price of the cloud's reach is distance, and distance is Latency. Use the simulator to feel it: switch from the “all edge” preset to the “all cloud” preset, then set single tasks to fog, and watch how much of the response time becomes network rather than computation.
- T1Image resizePT 8 ms
- T2Object detectionPT 40 ms
- T3Draw boxesPT 8 ms
Formula
RT = Σ one-way hops + Σ PT
RT = (5 + 15 + 15 + 5) + (8 + 40 + 8) = 96 ms
Client→Edge 5 · Edge→Fog 15 · Fog→Edge 15 · Edge→Client 5
Each hop costs the difference between the one-way latencies of the two layers: Device 0, Edge 5, Fog 20, Cloud 70 ms from the client.
Recall
Give the cloud's three “bigs” with an example of each.
Quick check
Which task belongs in the cloud layer rather than the fog?
Follow one temperature reading through a cold-storage company. The rule for frozen goods is that the room must stay at or below 8 °C, and one thermometer has just read 9 °C.
Worked example
One reading's journey up the tree
Sense
The thermometer, an end device, measures 9 °C and sends it to the warehouse gateway.Process locally
The Edge gateway compares the reading with the 8 °C threshold and switches on the backup cooler at once. No message has left the building yet.Coordinate regionally
The gateway forwards a short alarm, not the raw stream, to the regional fog node. The fog node sees alarms from every warehouse in the region and reroutes today's deliveries away from the warm one.Analyse globally
The fog node passes a daily summary to the cloud, which keeps a year of history from every region and learns which compressors are about to fail.Result
Each level acted on its own time scale and passed up something smaller than it received: a reading, then an alarm, then a summary.
The rule: sense, process, coordinate, analyse
Slide 13 compresses this into one sentence: edge devices sense, edge gateways provide local processing, fog nodes coordinate regionally, and the cloud analyses globally. In this course's model the devices and gateways belong to the Edge layer, the coordinators to the fog layer, and the analytics to the Cloud layer.
The shape is a tree. Many devices hang off each gateway, several gateways off each fog node, and a few fog nodes off the cloud. NIST describes the same structure when it says fog nodes can be clustered “vertically (to support isolation), horizontally (to support federation)”. Data flows up the tree and shrinks at every level, while decisions are taken at the lowest level that has enough information.
Bonomi's fog tiers from part 01 give the time scales, here in his Smart Grid example. The first fog tier runs control loops “from milliseconds to sub seconds”, and it “filters the data to be consumed locally, and sends the rest to the higher tiers”. Higher fog tiers work on “seconds to minutes”, and even days. The cloud keeps data for months to years.
Recall
State the one-line slogan for what each layer serves, and the four verbs of the IoT hierarchy.
Quick check
In the slide 13 hierarchy, which element hides the fog nodes from a smart thermostat?
A platoon of trucks drives at 108 km/h, which is 30 m/s, with gaps of about 10 m. The front truck brakes hard. How far does the next truck travel before it even starts to react? That depends on where the brake decision is computed.
Worked example
Metres travelled while waiting
Decision in the cloud
Assume a 60 ms round trip. d = 30 × 0.06 = 1.8 m, almost a fifth of the gap.Decision through a roadside unit
Assume a 5 ms round trip. d = 30 × 0.005 = 0.15 m.Decision on the car or over V2V
Assume about 1 ms. d = 30 × 0.001 = 0.03 m.Result
The roadside unit gives the truck about 1.65 m more margin than the cloud, and the car itself loses only centimetres. If both trucks brake equally hard, this waiting distance is exactly how much of the gap is lost by the time both stop. These latencies are illustrative assumptions, but the ratio is the point: distance in the network becomes distance on the road.
The rule: map each element to a layer by how fast it must act
- Cars are edge nodes. Each vehicle processes its own sensor data in real time, in the Edge layer. Cars also talk directly to each other over the vehicle-to-vehicle (V2V) network, which skips every piece of infrastructure.
- Roadside units are fog nodes. An Roadside unit (RSU) along the road sees the vehicles near it and can, for example, warn them about a traffic jam ahead. In the paper behind slide 14's figure, “RSUs are managed by the RSUC which has more resources for computing, storage, and communication through Internet to the cloud.” The RSU controller (RSUC) is therefore the fog aggregator that links the roadside units to the cloud.
- The cloud coordinates globally. The Cloud layer offers global coordination, long-term storage and HPC, but at a higher Latency.
The figure comes from Wang, Guo, Gao, Fang, Li and Sun, “Fog-Based Distributed Networked Control for Connected Autonomous Vehicles” (2020). In their connected cruise control case study the RSU calculates control coefficients and sends them to each vehicle, and the car's on-board controller computes the real-time control signals from them. The split mirrors the time scales: the fog tunes the controller, the car closes the loop.
This is not an exotic case. NIST uses it to illustrate the fog's geographical distribution: “streaming services to moving vehicles, through proxies and access points geographically positioned along highways”. Bonomi lists cars talking to roadside units and smart traffic lights as a core fog use case.
Recall
In the platoon, map each element to a layer and give one job for each.
Quick check
Why does a platoon's emergency brake warning travel car to car instead of through the cloud?
Recap
If you remember nothing else
- The fog gives compute, storage and networking close to clients, and serves many edges.
- Fog benefits: less edge-to-cloud traffic, geo-restricted data for compliance, and shorter edge-to-edge paths.
- The edge sits at the WAN boundary. It takes offloaded jobs, processes small data fast, and is the gateway that hides the fog.
- The cloud gives big compute (HPC), big storage (big data) and big networking (multi-cloud).
- Edge serves the client, fog serves many edges, cloud serves everyone.
- In an IoT tree, devices sense, gateways process locally, fog nodes coordinate regionally, and the cloud analyses globally.
- In a platoon, cars are edge nodes, RSUs and the RSUC are fog, and the cloud gives global coordination at higher latency.
- Sources draw the boundary differently: NIST puts gateways in the fog and end devices in the edge.
Sources
- NIST SP 500-325: Fog Computing Conceptual Model (Iorga et al., March 2018)PaperNISTFog and fog node definitions, six essential characteristics, Annex A on fog against edge, vertical and horizontal clustering.(opens in a new tab)
- Fog Computing and Its Role in the Internet of Things (Bonomi, Milito, Zhu, Addepalli, MCC 2012)PaperACMFog definition, Smart Grid tiers and time scales, fog localization and cloud globalization, traffic lights.(opens in a new tab)
- The Emergence of Edge Computing (Satyanarayanan, IEEE Computer, 2017)PaperIEEE Computer SocietyCloudlets reduce ingress bandwidth, the 100 Gbps and 8.5 Tbps video estimates, privacy policy, masking outages.(opens in a new tab)
- Fog-Based Distributed Networked Control for Connected Autonomous Vehicles (Wang et al., 2020)PaperWireless Communications and Mobile Computing, WileySource of slide 14's figure: RSUs, RSUC and connected cruise control. RSU/RSUC roles and the RSU-computed control coefficients are from Sections 3 and 4 of the full text.(opens in a new tab)
- RFC 9110: HTTP Semantics, §3.7 IntermediariesRFCIETFGateway (reverse proxy) definition, the parallel to an edge that hides the fog.(opens in a new tab)
- What is AWS IoT Greengrass?DocsAWSDevices act locally, filter and aggregate data, and react autonomously to local events.(opens in a new tab)
- Amazon S3DocsAWSObject storage for long-lived data at petabyte scale: the cloud's big storage.(opens in a new tab)
- High Performance Computing on AWSDocsAWSScaling HPC applications to thousands of CPUs and GPUs: the cloud's big compute.(opens in a new tab)
- Cross-Cloud Interconnect overviewDocsGoogle CloudDedicated links to other clouds at 10, 100 or 400 Gbps: the cloud's big networking.(opens in a new tab)
- Regulation (EU) 2016/679 (GDPR), Chapter V and Article 44DocsEUR-LexConditions for transferring personal data to third countries.(opens in a new tab)
- Standard Contractual Clauses for Personal Data TransferDocsSDAIAClauses for transferring personal data outside the Kingdom under the PDPL.(opens in a new tab)
- Amazon Alexa's new wake word research at InterspeechArticleAmazon ScienceOn-device wake word detection, cloud confirmation and recognition.(opens in a new tab)
- IEEE 1934-2018: OpenFog Reference ArchitectureDocsIEEEOptional further reading on fog architecture as a standard.(opens in a new tab)
Part 03: From legacy IT to virtualization
Starts the cloud computing section with the utility computing vision and the legacy IT stack, then defines virtualization, its history and its main forms.
6 concepts, slides 15-21
Why this part matters
This part opens the cloud half of the lecture. Every cloud product you will price in parts 05 and 06, classify in part 07 or deploy in parts 08 and 09 is, at heart, a way of handing some layers of the classic IT stack to someone else. Before you can say which layers move, you need to know what the stack is and which technology lets one provider sell slices of a machine to many customers at once.
That technology is virtualization, and it is what turned a 1961 economic idea into a business. For your research and for the edge systems you will design, the trade-offs between a virtual machine, an emulator and a container decide how much isolation you get, how much overhead you pay and how portable your workload is. This part builds those distinctions carefully, and it flags two places where the slides blur them.
By the end you can
- Explain McCarthy's utility computing vision and map it to modern cloud characteristics.
- List the layers of legacy IT and say who manages each one.
- Define a virtual machine and explain transparency using Popek and Goldberg's three properties.
- Distinguish virtualization, emulation and simulation with an example of each.
- Classify virtualization techniques on two independent axes: guest modification and hypervisor placement.
- Explain which resources are isolated and which are shared in the hardware virtualization stack.
Look at your electricity bill. You never bought a power station, you never hired the engineers who run it, and you can plug in a kettle or a server without asking anyone. You pay for the kilowatt-hours you actually used, and the grid is reached through a socket in the wall. Computing can be sold the same way.
The claim is older than most people expect. Speaking at MIT's centennial in 1961, John McCarthy predicted that computing may someday be organized as a public utility, just as the telephone system is. He imagined computer service companies with subscribers connected to them, where each subscriber pays only for the capacity actually used, yet has access to everything a very large system offers. That is Utility computing, and the slide's verdict is right: it is remarkably close to what we now call cloud computing.
Three parts of one sentence
Read McCarthy's sentence slowly and it splits into three ideas. First, there is shared central capacity: one very large system that many subscribers draw on. Second, usage is metered and billed, so you pay for what you consume rather than for what you own. Third, subscribers reach the system over a network, from wherever they are. Fifty years later, NIST SP 800-145 wrote the formal definition of cloud computing, and three of its five essential characteristics line up with these ideas almost word for word.
| McCarthy's words | What it means | NIST characteristic |
|---|---|---|
| Access to all programming languages characteristic of a very large system | Many subscribers share one big pool | Resource pooling |
| Pay only for the capacity actually used | Usage is metered and billed | Measured service |
| Subscribers connected to a computer service company | Reached over a network from anywhere | Broad network access |
Part 06 teaches all five NIST characteristics and the deployment models. For now, keep the pattern: cloud computing is first an economic arrangement, and the technology exists to make that arrangement safe and cheap.
Recall
Which three parts of McCarthy's 1961 vision line up with modern cloud characteristics?
Quick check
McCarthy's "pay only for the capacity that he actually uses" anticipates which NIST characteristic most directly?
Picture a company that runs its ERP system on premise. It leases a building and fits it with power and cooling, buys racks, switches, a storage array and servers, installs an operating system, runs a database, patches security holes and finally deploys the application its staff actually use. Every one of those steps needs its own people and its own budget line, and nobody else is responsible if any of them fails.
That is Legacy IT. The defining rule is simple: the customer manages every layer. The slide draws nine of them, and they are worth learning bottom to top, because every cloud service model in part 07 is defined by how many of these same layers move to a provider. Infrastructure as a service takes the bottom ones, platform as a service takes more, and software as a service takes them all.
Legacy IT layers, bottom to top, and what the customer must do
- Data center
- Provide the building, power, cooling and physical security that host applications and data.
- Networking
- Buy and run the switches, routers and cabling that physically connect the servers.
- Storage
- Provide disks and arrays that hold large amounts of data, plus backup.
- Server
- Buy, rack, repair and refresh the machines that run the business applications.
- Virtualization
- Optional: run a hypervisor to share servers. Classic legacy IT had none (see the note below).
- Operating system
- Install, license, patch and upgrade the OS on every machine.
- Databases
- Install, tune, back up and replicate the database engines.
- Security
- Run firewalls, access control, monitoring and patching across the stack.
- Applications
- Deploy, configure and maintain the software the users actually need.
The slide's four bullets name the expensive physical base of this stack: data center technology to operate applications and store data, servers to run the business applications, networking to connect the servers, and storage for large amounts of data. Those are exactly the layers that part 05 will price as capital expenditure, and they are the first layers a cloud provider takes over. NIST SP 800-145 phrases each service model the same way, by listing which layers the consumer still controls.
Recall
In legacy IT, which layers does the customer manage?
Quick check
In the legacy IT model, who manages the operating system layer?
Take a server with 64 cores and 256 GB of RAM that runs one web application at about 10% load. Most of that machine is idle most of the time, yet it still uses power, rack space and an administrator's attention. Now carve it into six virtual machines: a Linux web server, a Linux database, a Windows directory server and three more. Each one gets its own virtual CPUs, its own slice of RAM, its own virtual disk and its own network card, and each boots its own operating system.
That carving is Virtualization: the resources of one computer are split so that several independent operating system instances can run on it at the same time. Each piece is a Virtual machine (VM), and it has two defining traits. It behaves like any other computer, with its own components, and it runs inside an isolated environment on the physical machine. NIST SP 800-125 puts it formally: virtualization is the simulation of the software and hardware on which other software runs, and that simulated environment is called a virtual machine. The guest appears to have its own processor, memory, storage controllers, Ethernet controllers, display, keyboard and mouse.
Two payoffs: consolidation and isolation
The first payoff is consolidation. NIST calls it operational efficiency: you place more load on each physical computer, so you buy, power and cool fewer of them. The hardware stops being a set of fixed boxes and becomes a pool, which is exactly the Resource pooling that McCarthy's utility needed.
The second payoff is isolation. Smith and Nair state it plainly: if security on one guest system is compromised, or one guest operating system fails, the software on the other guests is not affected. NIST adds that the hypervisor partitions the system's resources so that each guest can reach only its own. Isolation is what lets a provider rent neighbouring VMs to strangers, and it is the "safely" in the slide's summary that virtualization turns hardware into a pool that can safely run many independent computers.
Worked example
Consolidating five lightly loaded servers (illustrative numbers)
State the starting point
Five physical servers of identical size run at 10%, 12%, 15%, 11% and 14% average CPU utilization. These figures are made up for teaching, but single-digit to low-teen utilization is the typical shape of one-application-per-server estates.Add the loads as if they shared one host
10 + 12 + 15 + 11 + 14 = 62% of one server of the same size.Leave headroom
The combined load is well under 100%. The remaining 38% absorbs peaks that do not coincide and the small overhead of the hypervisor itself.Result
One host replaces five, each workload keeps its own OS and isolation, and four machines' worth of power, space and maintenance disappears. If peaks did line up, you would size for the peak instead of the average.
Recall
What are the two payoffs of virtualization, and what does each one buy a cloud provider?
Install Ubuntu in VirtualBox on an x86 laptop. Ubuntu's own code runs directly on the real CPU at full speed. Only when Ubuntu tries something sensitive, such as changing page tables or touching a device, does the virtualization software step in, perform the operation on Ubuntu's behalf against the real hardware, and hand back the result. Ubuntu never notices. Now run an old console game in a Nintendo emulator: every guest instruction has to be translated, because the guest speaks a different instruction set. (BlueStacks, the slide's other example, does the same for the ARM code inside Android apps.) Finally, sit in a flight simulator. It does not run the aircraft's software at all; it models how the aircraft behaves.
These three scenes are virtualization, emulation and Simulation. The first property to name is transparency. Inside a VM you install an operating system and applications exactly as on a physical computer, and the applications do not notice they are in a VM. Requests from the guest are transparently intercepted by the Virtual machine monitor (VMM) and converted for the real hardware, and the VM never becomes aware of the layer between it and the machine. Smith and Nair describe the mechanism: when a guest performs a privileged instruction, the VMM intercepts it, checks it and performs it on behalf of the guest, and guest software is unaware of this behind-the-scenes work.
Popek and Goldberg: what counts as a VM
In 1974, Popek and Goldberg gave the precise version. A virtual machine is "an efficient, isolated duplicate of the real machine", and a VMM must have three properties.
- Equivalence. Programs see an environment essentially identical to the original machine. Apart from timing and resource availability, they behave as they would on bare hardware.
- Efficiency. A statistically dominant subset of the guest's instructions runs directly on the real processor, with no software intervention.
- Resource control. The VMM stays in complete control of the real resources. A guest cannot reach resources it was not allocated.
The efficiency property is the one that draws the line. Popek and Goldberg say explicitly that it rules out traditional emulators and complete software interpreters (simulators) from the virtual machine umbrella. Their Theorem 1 gives a sufficient condition for building a VMM this way: every sensitive instruction is also privileged, so it traps to the VMM instead of running silently.
| Aspect | Virtualization | Emulation | Simulation |
|---|---|---|---|
| Core idea | Creates logical copies of physical resources | Mimics a hardware/software interface | Models a system's behaviour |
| What runs natively | Most guest instructions | Nothing: every instruction is translated | The model, not the real software |
| Same ISA required | Yes | No, that is the point | Not applicable |
| Guest aware? | No, interception is transparent | No, the translation is hidden | There is no guest |
| Speed | Close to native | Much slower than native | Whatever the model needs |
| Use case | Running many OSs or VMs on one host | Running software built for other hardware | Training or testing |
| Example | VMware, VirtualBox | BlueStacks, Nintendo emulator | Flight simulator, VR training |
Emulation still matters. Smith and Nair note that a whole-system VM whose ISA differs from the host, such as Virtual PC running Windows on a PowerPC Mac, must emulate both application and OS code. It works, and it is how you run old consoles or foreign architectures, but you pay for every instruction.
Recall
State Popek and Goldberg's three VMM properties, and the one that excludes emulators.
Quick check
Software interprets almost every guest instruction because the guest uses a different ISA. By Popek and Goldberg's definition, what is it?
In 1972 an IBM mainframe could be shared by hundreds of users at once. The trick was the Control Program (CP), which gave each user a complete virtual machine, a virtual System/370 duplicating the physical hardware. Inside that VM each user ran CMS, a simple single-user operating system. Many single-user machines, each in its own VM, added up to a multi-user time-sharing system. That product was VM/370, and Creasy's 1981 paper cited on the slide tells its story.
The lineage starts earlier. CP/CMS was conceived in 1964, CP-40 and CMS were running in 1966, and CP-67 on the System/360 Model 67 was in production use from 1967. IBM announced VM/370 on August 2, 1972. Its descendant, z/VM, still runs on IBM mainframes, so virtualization is one of the longest-lived ideas in computing.
A family, classified by what is virtualized
Today "virtualization" names a family of techniques. They differ in which layer is virtualized and how much of the guest is changed.
- Partitioning splits one machine's hardware into fixed, separate partitions.
- Hardware emulation presents hardware different from the real platform, at a translation cost.
- Application virtualization wraps one application so it runs isolated from the host OS.
- Full virtualization presents virtual hardware close enough to the real thing that unmodified guest OSs run.
- Paravirtualization modifies the guest kernel to call the hypervisor directly instead of touching emulated hardware.
- Hardware virtualization gives each complete guest OS its own virtual hardware, which the VMM maps onto shared physical CPU, memory, storage and network (slide 21). CPU extensions such as Intel VT-x andAMD-V (hardware-assisted virtualization) make this efficient.
- OS-level virtualization shares one kernel and isolates user spaces, which gives containers.
- Storage virtualization pools disks behind a SAN or software-defined storage.
- Network virtualization carves logical networks out of shared links with VLAN or VXLAN.
Hardware assistance exists because x86 does not meet Theorem 1's condition. Barham and colleagues note that some x86 supervisor instructions fail silently instead of trapping, so a VMM cannot catch them. VMware ESX Server solved this for full virtualization by dynamically rewriting guest code, Xen sidestepped it with paravirtualization, and later CPUs added extensions that make the instructions trap.
Two independent axes
The most important correction in this part is that two different questions get mixed up. The first axis is guest modification: does the guest OS run unmodified (full virtualization), or is its kernel changed to cooperate with the hypervisor (paravirtualization)? Barham and colleagues define paravirtualization exactly this way: a virtual machine abstraction similar but not identical to the hardware, which requires modifications to the guest OS but none to guest applications.
The second axis is hypervisor placement: does the VMM run directly on the hardware (Type 1 hypervisor) or on top of a host operating system (Type 2 hypervisor)? Goldberg's 1973 thesis defines it so: a Type I VMM runs on a hardware host, a Type II VMM runs on an extended host, meaning an operating system. Part 04 teaches the two types in depth. NIST SP 800-125 confirms the axes are independent by naming bare metal and hosted as the two forms of full virtualization.
| Unmodified guest (full) | Modified guest (para) | |
|---|---|---|
| Type 1, bare metal | VMware ESX/ESXi, Xen HVM | Xen PV |
| Type 2, hosted | VirtualBox, VMware Workstation | Paravirtual storage and network drivers in hosted products |
Recall
Why is "paravirtualization = Type 1" wrong?
Quick check
Which statement correctly fixes slide 20's mapping of virtualization techniques to hypervisor types?
Three VMs share one host. Follow a single disk write from one of them. The application calls a library function to save a file. The library asks the guest OS. The guest OS drives what it believes is its disk controller, but that controller is virtual hardware. When the guest touches it, the access reaches the hardware/software interface, where the Virtual machine monitor (VMM) catches it and maps it onto the real SSD that all three VMs share.
Saves a file.
Calls the OS.
Drives its disk driver.
A virtual disk controller.
Maps virtual to physical.
CPU, memory, storage, network.
The rule behind the picture is a single horizontal cut. Everything above the hardware/software interface, the application, library, operating system and virtual hardware, is replicated for each VM and isolated from the others. Everything below it, the CPU, memory, storage and network, is physical, shared and multiplexed by the VMM. Smith and Nair say the VMM has access to, and manages, all the hardware resources, and NIST says the hypervisor controls the flow of instructions between the guests and the physical CPU, disk, memory and network cards. So the OS instances are isolated from each other, and each runs on virtualized resources that are mapped to physical ones.
Worked example
Tracing one disk write through the layers
Application and library
The program calls a file write. The C library turns it into a system call to the guest kernel. Nothing special has happened yet: this all runs natively on the real CPU.Guest OS
The guest kernel's file system and disk driver build a request and write it to the registers of its disk controller, exactly as they would on a physical machine.Interception at the interface
Those controller registers are virtual. The access traps to the VMM, which checks that the request touches only this VM's virtual disk.Shared hardware
The VMM translates the virtual block address to a location in the VM's disk image on the shared SSD, issues the real I/O, and later signals completion back into the guest as a virtual interrupt.Result
The guest saw an ordinary disk write. Isolation held because the VMM only ever mapped it onto this VM's share of the device.
Where you cut decides the kind of virtualization
The interface on the slide is essentially the instruction set architecture. Smith and Nair write that the ISA marks the division between hardware and software, and they identify three key interfaces: the ISA, the ABI (system calls plus user instructions) and the API (library calls). Cut the stack at the ISA and you get a system VM with a full guest OS, as on this slide. Cut higher, at the ABI or API, and you get process VMs or OS-level containers, where the kernel is shared and only the user space is isolated. That is the bridge to part 04.
Recall
In the hardware virtualization stack, what is isolated and what is shared?
Recap
If you remember nothing else
- McCarthy (1961) described computing sold like a utility: shared capacity, pay per use, network access. That is the economic idea behind cloud computing.
- In legacy IT the customer runs every layer, from the data center to applications. Cloud models move layers to a provider.
- A VM is an efficient, isolated duplicate of a real machine. The guest OS does not notice the VMM intercepting sensitive operations.
- Virtualization runs most instructions natively on the same ISA. Emulation translates between ISAs. Simulation models behaviour without running the real software.
- Virtualization began with IBM CP-40 and CP-67, and VM/370 (announced 1972) turned time-sharing into one VM per user.
- Full versus para (is the guest modified?) and Type 1 versus Type 2 (where does the hypervisor run?) are independent axes.
- Above the hardware/software interface everything is isolated per VM. Below it, the physical hardware is shared.
Sources
- Formal Requirements for Virtualizable Third Generation ArchitecturesPaperPopek and Goldberg, Communications of the ACM 17(7), 1974Equivalence, efficiency, resource control, and Theorem 1.(opens in a new tab)
- The Origin of the VM/370 Time-sharing SystemPaperCreasy, IBM Journal of Research and Development 25(5), 1981CP-40, CP-67 and VM/370 history.(opens in a new tab)
- Xen and the Art of VirtualizationPaperBarham et al., SOSP 2003Definition of paravirtualization and the x86 trap problem.(opens in a new tab)
- The Architecture of Virtual MachinesPaperSmith and Nair, IEEE Computer 38(5), 2005ISA, ABI and API interfaces; isolation; process versus system VMs.(opens in a new tab)
- Architectural Principles for Virtual Computer SystemsPaperGoldberg, Harvard PhD thesis, 1973Origin of the Type I and Type II VMM distinction.(opens in a new tab)
- SP 800-125: Guide to Security for Full Virtualization TechnologiesDocsNIST, 2011Definition of a VM, bare metal and hosted full virtualization, hardware emulation.(opens in a new tab)
- SP 800-145: The NIST Definition of Cloud ComputingDocsNIST, 2011(opens in a new tab)
- z/VM History: TimelineDocsIBMVM/370 announced August 2, 1972.(opens in a new tab)
- The Cloud ImperativeArticleGarfinkel, MIT Technology Review, 2011McCarthy's 1961 utility computing quote.(opens in a new tab)
- Architects of the Information SocietyBookGarfinkel, MIT Press, 1999(opens in a new tab)
- Understanding Full Virtualization, Paravirtualization, and Hardware AssistArticleVMware, 2007Further reading.(opens in a new tab)
Part 04: Hypervisors, containers, namespaces and cgroups
Compares Type 1 and Type 2 hypervisors, then explains OS-level virtualization (containers) built from Linux namespaces and control groups.
6 concepts, slides 22-28
Why this part matters
Every cloud service in the rest of this course, from EC2 and Lambda to Kubernetes on EKS or GKE and edge nodes running k3s, is built from one of two isolation designs. Either a hypervisor splits the hardware into virtual machines, or one shared kernel is split into containers by namespaces and cgroups.
The choice between them decides overhead, density, start-up time and how strong the security boundary is. Comparing Type 1, Type 2 and containers, and telling namespaces from cgroups, are the likely exam questions. The same choice comes back when you size edge devices for the research project, where a small gateway may not have the memory to boot a full guest OS for every service.
By the end you can
- Place Type 1 and Type 2 hypervisors in the software stack and name real examples of each, including where KVM and Hyper-V fit.
- Compare Type 1 and Type 2 on overhead, resource management, administration and isolation, and explain why cloud providers choose Type 1.
- Explain OS-level virtualization: what containers share, what they isolate, and why their security boundary is weaker than a VM's.
- Distinguish namespaces (what a process sees) from cgroups (what a process uses), and name the main namespace types and cgroup functions.
- Read a cgroup CPU quota and weight and predict how a container behaves under contention and at its memory limit.
Start with a laptop running VirtualBox with Ubuntu inside it. When you press the power button, macOS boots first. VirtualBox is then just another application, next to Safari, and the whole Ubuntu guest lives inside that application's process. Now compare a server in an AWS data center that hosts EC2 instances. Nothing general purpose boots before the hypervisor. It is the first thing on the machine, and it hands out CPU and memory to the guests directly.
Both machines do Virtualization, and both run a Virtual machine monitor (VMM) that makes each Virtual machine (VM) believe it has a real computer. The difference is where the VMM sits, and that comes down to one question: who owns the hardware?
Two placements
- A Type 1 hypervisor (bare metal, or native) is the lowest software layer on the machine. It runs directly on the hardware and acts as a specialized OS whose only job is to run guests. Red Hat describes it as a hypervisor that "runs directly on the host's hardware to manage guest operating systems".
- A Type 2 hypervisor (hosted) runs on a conventional OS "as a software layer or application". The host OS owns the hardware, and the VMM is one of its processes, with less privilege than the OS itself. VirtualBox's own manual calls it "a so-called hosted hypervisor, sometimes referred to as a type 2 hypervisor".
Guest OS 1 to N run directly on the hypervisor, which owns CPU, memory, storage and network.
The VMM is an app on the host OS. Each guest request crosses the VMM and the host kernel.
Real examples, and where the labels get blurry
VMware's Type 1 hypervisor is ESXi, described as "the hypervisor [that] runs virtual machines"; vCenter Server is the central management service for many ESXi hosts, and vSphere is the name of the whole suite. On the hosted side sit VirtualBox, VMware Workstation and Parallels Desktop.
Picture a lab PC running VirtualBox for a coursework VM, next to an EC2 c7i instance. When the coursework VM compiles code, its virtual CPU is just a thread that competes with Chrome and the antivirus for time on the host scheduler. The EC2 instance gets CPUs and memory partitioned for it by a minimal hypervisor, with no desktop OS in the way.
That picture gives the whole comparison. Because a Type 1 hypervisor owns the hardware, it adds less overhead (the slide says about 5%), manages resources directly and isolates guests more strongly, but it needs more skill to install and operate. A Type 2 hypervisor reaches the hardware only through the host OS, so it costs more (about 10%), controls resources indirectly and isolates less, but anyone can install it.
| Property | Type 1 (bare metal) | Type 2 (hosted) |
|---|---|---|
| Runs on | Directly on the hardware, as the lowest software layer | As a process of a general purpose host OS |
| Approx. overhead | ≈5% | ≈10% |
| Resource management | Direct: the hypervisor schedules CPUs and partitions memory itself | Indirect: every request goes through the host OS scheduler and drivers |
| Administration | Needs more skill: a dedicated server, remote management, clusters | Easy: install it like any desktop app |
| Isolation | Stronger: only a small hypervisor sits between tenants | Weaker: a compromise of the host OS exposes every guest |
| Typical users | Cloud providers and enterprise data centers | Developers, students, testers on laptops |
| Examples | VMware ESXi, Microsoft Hyper-V, KVM (debated), AWS Nitro | Oracle VirtualBox, VMware Workstation, Parallels Desktop |
Why cloud providers pick Type 1
A Cloud service provider (CSP) sells slices of the same physical server to strangers. Every percent of overhead is capacity it cannot sell, and every extra layer between tenants is attack surface. AWS shows where this leads. Its Nitro Hypervisor is "a custom-developed, minimized hypervisor based on KVM", cut down on purpose: it has "no networking stack, no general-purpose file system implementations, and no peripheral device driver support". Networking and storage are offloaded to dedicated Nitro Cards and exposed to guests through SR-IOV. The hypervisor does little more than partition CPU and memory, which is the Type 1 idea taken to its limit, and the design is lean enough that AWS can also sell bare metal instances.
Take the slide's own example: one server runs a database and a web server as plain processes. Two things can go wrong. First, the web server can list the database's processes, send them signals, read its files and bind its ports, so whoever compromises one has a view of the other. Second, a heavy query makes the database use every CPU core, and the website stalls.
These are different problems, and Linux fixes each with a different kernel feature. The first is about visibility. A Namespace gives a group of processes its own view of a kernel resource, so the web server cannot even see the database's processes, network stack or mount points. The second is about consumption. Control groups (cgroups) put quotas, weights and accounting on how much CPU, memory and I/O a group can use. Docker's security guide gives the same split: namespaces "provide the first and most straightforward form of isolation", while cgroups "implement resource accounting and limiting".
| Problem | Fix | Kernel feature | Question it answers |
|---|---|---|---|
| Processes see and touch each other (PIDs, ports, files) | Give each group its own view of kernel resources | Namespaces | What can I see? |
| Processes compete for the same CPU, memory and I/O | Put quotas, weights and accounting on resource use | Control groups (cgroups) | How much can I use? |
Neither feature is new. The first namespace (mount) arrived in Linux 2.4.19 in 2002, most of the rest landed between 2.6.15 and 2.6.26 (July 2008), and user namespaces were only completed in 3.8 (2013), which made unprivileged containers practical. Cgroups were started in 2006 and merged in 2.6.24. A Container is simply these two combined, plus a root filesystem image. NIST describes the second half with a concrete case: a host with 10 GB of memory can give 1 GB to each of nine containers, so no single one can starve the rest.
You can make a Namespace by hand on any Linux machine. The command below starts a shell in a new PID namespace and mounts a fresh /proc for it. Inside, ps a shows only two processes, and the shell believes it is PID 1.
sudo unshare --fork --pid --mount-proc bash
ps aPID TTY STAT TIME COMMAND
1 pts/0 S 0:00 bash
8 pts/0 R+ 0:00 ps aFrom a second terminal on the host, the same bash appears in ps -A under a large global PID such as 24517. Nothing was copied and no second kernel was started. The kernel simply shows that process a different numbering. NIST gives the container version: ps -A inside an Apache Container lists only httpd.
The rule
The Linux manual defines a namespace as something that "wraps a global system resource in an abstraction that makes it appear to the processes within the namespace that they have their own isolated instance". There is one kernel and many views. Every process belongs to exactly one namespace of each type, and a container runtime such as Docker creates a fresh set (all but user and time, by default) for every container. You can see a process's memberships as symbolic links in /proc/<pid>/ns/ (since Linux 3.8), and a tool can join an existing namespace with setns(2), which is how docker exec gets a shell into a running container.
| Namespace | Flag | What it isolates |
|---|---|---|
| Cgroup | CLONE_NEWCGROUP | The cgroup root directory a process sees |
| IPC | CLONE_NEWIPC | System V IPC objects and POSIX message queues |
| Network | CLONE_NEWNET | Network devices, IP stacks, routing tables, ports |
| Mount | CLONE_NEWNS | Mount points, so each container has its own filesystem tree |
| PID | CLONE_NEWPID | Process ID numbers, starting again at 1 |
| Time | CLONE_NEWTIME | Boot and monotonic clock offsets |
| User | CLONE_NEWUSER | User and group IDs, so root inside can be unprivileged outside |
| UTS | CLONE_NEWUTS | Hostname and NIS domain name |
The network namespace is what gives each container "its own networking stack, including unique interfaces and IP addresses", in NIST's words, so two containers can both listen on port 80. The user namespace, when enabled, lets a process be root inside its container while mapping to an unprivileged user on the host.
On a host with 2 CPUs, start MongoDB with docker run --cpus=1.5 --memory=512m mongo. Docker writes two numbers into the container's cgroup: a CPU quota of 150000 µs per 100000 µs period, and a memory ceiling of 512 MiB. If MongoDB grows past that ceiling and the kernel cannot reclaim memory, the OOM killer kills a process inside that cgroup only. The web server next to it never notices.
docker run -d --name db --cpus=1.5 --memory=512m mongo
cat /sys/fs/cgroup/system.slice/docker-$(docker inspect -f '{{.Id}}' db).scope/cpu.max
cat /sys/fs/cgroup/system.slice/docker-$(docker inspect -f '{{.Id}}' db).scope/memory.max150000 100000
536870912How cgroups work
The manual says cgroups "allow processes to be organized into hierarchical groups whose usage of various types of resources can then be limited and monitored". Each kind of resource is enforced by a controller, which Red Hat defines as "a kernel subsystem that represents a single resource, such as CPU time, memory, network bandwidth or disk I/O". The main controllers are cpu, memory, io, pids and cpuset; freezing, a separate freezer controller in cgroup v1, is the core file cgroup.freeze in v2. Cgroup v2, official since Linux 4.5, puts everything in one tree mounted at /sys/fs/cgroup, where every setting is a small file you can read and write.
The slide lists four functions, and each maps to concrete files:
The four cgroup functions and the cgroup v2 files behind them
- Limiting
- Hard caps: cpu.max, memory.max, io.max, pids.max
- Prioritization
- Shares under contention: cpu.weight, io.weight
- Accounting
- Measured use for monitoring and billing: cpu.stat, memory.current, io.stat
- Control
- Freezing and thawing a whole group: cgroup.freeze (the basis for checkpoint and restart with CRIU)
Worked example
Reading a CPU quota and a CPU weight
Turn cpu.max into CPUs
The file holds 150000 100000: a quota of CPU time, then the period it applies to. The format is $MAX $PERIOD, and the default is max 100000 (no limit).
Compare with the machine
On a 2-CPU host that is 1.5 / 2 = 75% of all CPU time. It is a hard ceiling: once the group has used 150 ms of CPU time in a 100 ms period, it waits for the next period, even if the machine is otherwise idle.
Add prioritization
Now two busy containers have --cpu-shares 1024 and 512 (cgroup v2 calls this cpu.weight, default 100, range 1 to 10000). When they compete, each gets its weight divided by the sum of weights:
This is a soft limit. If the second container goes idle, the first may use all the CPU it can get.
Result
Limits are ceilings that hold even on an idle machine; weights are shares that matter only under contention. That is how to read the slide's diagram where four groups each get 25%: with equal weights, each gets a quarter only while all four are busy.
Recall
In one sentence each, where does a Type 1 and a Type 2 hypervisor run, and who owns the hardware?
Recall
Why do cloud providers choose Type 1 hypervisors?
Recall
What is shared and what is isolated in OS-level virtualization?
Recall
Namespaces and cgroups: which question does each answer?
Recall
A container gets --cpus=1.5. What does cpu.max hold, and what does it mean?
Recall
Two busy containers have cpu.weight 200 and 100 and no cpu.max. What share does each get, and what happens when the second goes idle?
Quick check
A cloud provider wants minimal overhead and strong isolation between tenant VMs on each server. Which virtualization layer should it deploy?
Quick check
A database and a web server share one Linux host. The database starts using every CPU core and the website slows down. Which kernel feature fixes this?
Quick check
Inside a container, ps shows only two processes and your app has PID 1. What produces this view?
Quick check
Why can a Linux host run a Windows VM but not a native Windows container?
Recap
If you remember nothing else
- Type 1 hypervisors run on bare metal and act as the OS (ESXi, Hyper-V, KVM). Type 2 run as an app on a host OS (VirtualBox, VMware Workstation, Parallels).
- Type 1: about 5% overhead, direct resource control, stronger isolation, harder to run. Type 2: about 10% overhead, indirect control, easier, less isolation.
- Cloud providers use Type 1. AWS Nitro is a minimal KVM-based Type 1 hypervisor that offloads I/O to dedicated cards.
- Containers use no hypervisor. They share the host kernel below the system call interface and isolate only their user space.
- Namespaces control what a process can see: PID, network, mount, UTS, IPC, user, cgroup and time.
- Cgroups control what a process can use: limiting, prioritization, accounting and control (freezing a whole group, which CRIU builds on to checkpoint and restart).
- A container is several namespaces plus cgroups plus a root filesystem. In the cloud, containers usually run inside VMs.
Sources
- What is a hypervisor?DocsRed HatType 1 and Type 2 definitions and examples(opens in a new tab)
- What is KVM?DocsRed HatEach VM is a regular Linux process(opens in a new tab)
- Kernel Virtual MachineDocsKVM project(opens in a new tab)
- The Definitive KVM API DocumentationDocsLinux kernel(opens in a new tab)
- Hyper-V ArchitectureDocsMicrosoft LearnRoot and child partitions, VMBus, enlightened I/O(opens in a new tab)
- vSphere Software ComponentsDocsBroadcom TechDocsESXi is the hypervisor; vSphere is the suite(opens in a new tab)
- Oracle VirtualBox manual: IntroductionDocsOracle(opens in a new tab)
- Security Design of the AWS Nitro System: the componentsDocsAmazon Web Services(opens in a new tab)
- Security Design of the AWS Nitro System: the Nitro System journeyDocsAmazon Web Services(opens in a new tab)
- SP 800-190: Application Container Security GuidePaperNIST (Souppaya, Morello, Scarfone, 2017)Sections 2.2 and 3.5.2(opens in a new tab)
- namespaces(7)DocsLinux manual pages(opens in a new tab)
- unshare(1)DocsLinux manual pages(opens in a new tab)
- cgroups(7)DocsLinux manual pages(opens in a new tab)
- Control Group v2DocsLinux kernelcpu.max, cpu.weight, memory.max, io.max, cgroup.freeze(opens in a new tab)
- Freezer SubsystemDocsLinux kernel(opens in a new tab)
- Resource constraintsDocsDocker(opens in a new tab)
- Docker Engine securityDocsDocker(opens in a new tab)
- Setting limits for applicationsDocsRed Hat (RHEL 8)(opens in a new tab)
- CRIU: Checkpoint/Restore In UserspaceDocsCRIU project(opens in a new tab)
- Firecracker: Lightweight Virtualization for Serverless ApplicationsPaperUSENIX NSDI 2020 (Agache et al.)(opens in a new tab)
- Performance Overhead Comparison between Hypervisor and Container based VirtualizationPaperIEEE AINA 2017 (Li, Kihl, Lu, Andersson)(opens in a new tab)
- An updated performance comparison of virtual machines and Linux containersPaperIEEE ISPASS 2015 (Felter et al.)Further reading on VM and container performance(opens in a new tab)
Part 05: Cost of on-premise operation and TCO
Lists the cost drivers of legacy IT and works through a three-year total cost of ownership calculation for a 50-employee news streaming SME.
5 concepts, slides 29-34
Why this part matters
Whether to move to the cloud is a cost decision before it is a technical one. To say “the cloud is cheaper” honestly, you first need a defensible number for what running your own servers really costs.
This part builds that baseline for a realistically sized company: a 50-person news streaming firm with five servers. You will meet the five places money goes in legacy IT, the accounting line between buying assets and paying to run them, and then the full arithmetic that lands on a three-year total cost of ownership of 89,715 EUR. Recomputing these numbers is a likely exam question. The method also carries straight into research, where you will size edge and fog deployments the same way. Part 06 takes this baseline and puts the cloud next to it.
By the end you can
- List the five cost drivers of legacy IT and classify each as one-off or recurring.
- Distinguish CapEx from OpEx using accounting definitions, and correct the slide's placement of maintenance.
- Recompute the case study's CapEx, annual OpEx and three-year TCO from the hardware and price inputs.
- Convert measured server power and cooling into annual energy cost and an implied PUE.
- Critique the TCO's assumptions and name the costs it leaves out.
A 50-person company buys five servers. The invoice arrives and gets paid, and it feels like the job is done. It is only the start. The servers need a rack and a room to stand in. They burn electricity every hour of every day, and an air conditioner has to remove the heat they make. Someone has to install them, configure them and patch them. In year three a disk dies, the RAM is too small, and the cycle of buying begins again.
That is the life of Legacy IT. In part 03 you saw that in this model the customer manages every layer, from the building up to the application. The flip side is simple: the customer also pays for every layer. The honest way to price it is the Total cost of ownership (TCO), which adds up everything it costs to own and run the system over its life, not just what it costs to buy.
Five buckets of cost
- Hardware purchase. Servers, storage and network gear.
- Housing. The enclosure (rack, power strips, console switch) and the room it lives in.
- Operations. Power, cooling and maintenance.
- Personnel. The people who install, configure and maintain the machines.
- Upgrades and refresh. Extending or replacing hardware as it ages.
The buckets behave differently over time. Hardware, housing kit and refreshes are lumpy: big amounts at a few moments. Accountants call this kind of spending Capital expenditure (CapEx). Operations and personnel recur month after month, and that is Operational expenditure (OpEx). The next concepts make this split precise. For now, notice which buckets the case study will actually price, because the gap matters at the end.
| Cost bucket | Example | One-off or recurring | Counted in the case study? |
|---|---|---|---|
| Hardware purchase | Servers, SAN, network switches | One-off | Yes |
| Housing | Rack, PDU, KVM switch, the room and its rent | Both: kit is one-off, rent recurs | Yes |
| Operations | Power, cooling, maintenance | Recurring | Power and cooling yes, maintenance no |
| Personnel | Administrators who install, configure and patch | Recurring | No |
| Upgrades and refresh | New disks, more RAM, replacement servers | Periodic | No |
Recall
Name the five cost buckets of legacy IT, and say which ones recur.
Hardware purchase, housing, operations (power, cooling, maintenance), personnel, and upgrades or refresh. Operations and personnel recur. Hardware and refresh are lumpy one-off or periodic spending, and housing is both: kit bought once plus rent every month.
The company in the case study is a small news streaming firm with 50 employees. It needs three things from its IT: networking, web hosting and backups. The question is what it costs to run that on its own hardware for three years, so it can later be set against a cloud bill for the same job.
Fix the assumptions before you add anything up
A Total cost of ownership (TCO) comparison is only fair when both scenarios use the same scope and the same assumptions. Otherwise the cheaper option may simply be covering less. The case study fixes three assumptions up front.
- The software is open source, so licence fees are zero on both sides.
- Maintenance costs the same on-premise and in the cloud, so it cancels out of the comparison and is left out of both totals.
- The horizon is 3 years. That matches how long servers are typically used before replacement. Barroso et al. note that servers have a shorter lifetime than buildings and are usually depreciated over 3 to 4 years, while data center buildings are depreciated over 15to 20 years.
The hardware
The firm buys five identical 1U rack servers. Each has two Intel Xeon E5-2640 v2 processors, a 2013 server part that Intel lists with 8 cores, 16 threads, a 2.0 GHz base clock, a 20 MB cache and a 95 W thermal design power, for two-socket boards only.
Server specification from the case study (per server unless stated)
- Quantity
- 5 servers, 1U each
- CPU
- 2 × Intel Xeon E5-2640 v2
- Per CPU
- 8 cores, 16 threads, 2.0 GHz base, 20 MB cache, 95 W TDP
- Memory
- 16 GB RAM
- Network
- 4 NICs × 4 ports
- Storage
- 5 TB SAS
- Power supply rating
- 460 W
- Measured draw (slide 34)
- 308 W
Two quick derivations make this concrete. Compute first: 5 × 2 × 8 = 80 physical cores, or 160 hardware threads, across the fleet. Power second: the two CPUs alone can dissipate 2 × 95 = 190 W, the power supply is rated for 460 W, and the measured draw used later is 308 W, about 67% of the rating.
Recall
Why does the case study leave maintenance cost out?
It is assumed to be the same in the on-premise and cloud scenarios, so it cancels when the two are compared. Leaving it out keeps the comparison fair, but it also means neither total is a full cost.
The five servers cost 17,500 EUR. They are bought once and used for years. The electricity bill arrives every month. A technician who comes to replace a failed fan is paid for that visit, even though the fan sits inside a long-lived asset. Those three payments fall into two different accounting categories, and the difference is the heart of the cloud's financial argument.
The accounting definitions
Capital expenditure (CapEx) is spending on fixed assets that will be used for more than one period. The international standard IAS 16 defines such property, plant and equipment as tangible items “expected to be used during more than one period”. Their cost is not charged all at once. Instead it is spread over the useful life by depreciation, “the systematic allocation of the depreciable amount of an asset over its useful life”. Barroso et al. put it in data center terms: CapEx is the investment made upfront and then depreciated over a certain timeframe.
Operational expenditure (OpEx) is the recurring cost of actually running things, excluding depreciation: electricity, repairs and maintenance, salaries of on-site staff, rent, and service subscriptions such as Software as a Service (SaaS), Platform as a Service (PaaS) and Infrastructure as a Service (IaaS). IAS 16 paragraph 12 is explicit that the costs of the day-to-day servicing of an asset are not added to its carrying amount. They are expensed as they occur.
| Aspect | CapEx | OpEx |
|---|---|---|
| When paid | Upfront and lumpy | Recurring and smooth |
| Accounting | Capitalised, then depreciated over the useful life | Expensed in the period it is incurred |
| Examples | Servers, SAN, switches, racks | Power, cooling, rent, salaries, repairs, subscriptions |
| Cloud analogue | None: the provider owns and depreciates the hardware | Pay-as-you-go instances; reserved commitments are prepaid OpEx |
| Risk | Over- or under-provisioning is locked in | Spending scales with actual use |
Why the cloud talks about this split
The Berkeley “Above the Clouds” report notes that the cloud's appeal is often described as converting capital expenses into operating expenses, but argues that “pay as you go” captures the real benefit better. Without an upfront purchase, money stays free for the core business, and capacity can follow demand instead of being bought for a guessed peak. AWS's Well-Architected cost guidance says the same in practical terms: adopt a consumption model, and stop spending money on the undifferentiated heavy lifting of racking, stacking and powering servers. That argument is exactly what part 06 tests against this case study.
Quick check
Under standard accounting, where does routine maintenance of a server belong?
Recall
Is routine server maintenance CapEx or OpEx, and why?
OpEx. Day-to-day servicing does not create or extend an asset, so it is expensed as incurred (IAS 16paragraph 12).
Now the firm goes shopping. Everything on this list is a durable asset paid for on day one, so all of it is Capital expenditure (CapEx). The rule is a plain sum: quantity times unit price, over every item bought.
Worked example
Summing the CapEx
Servers
5 × 3,500 EUR = 17,500 EUR.
Storage area network
One SAN with 5 TB of shared storage: 35,000 EUR.
Network switches
4 × 3,677.50 EUR = 14,710 EUR.
Facilities
Power distribution unit and KVM switch for one rack: 897 EUR.
Cooling equipment
For one rack: 717 EUR.
Result
17,500 + 35,000 + 14,710 + 897 + 717 = 68,824 EUR of CapEx, paid before a single request is served.
Storage, not compute, dominates
The surprise in the bill is where the money goes. The servers, the part everyone pictures, are only about a quarter of it. The SAN alone is 35,000 / 68,824 ≈ 50.9% of the CapEx, and once running costs are added it is still 39.0% of the whole three-year Total cost of ownership (TCO). A SAN is a dedicated network that gives servers shared, redundant, block-level access to pooled disks, and that redundancy and shared access are expensive. This is one reason managed storage and Object storage are such strong selling points for the cloud: the provider amortises that cost over thousands of customers.
Decoding the small items
- PDU, a power distribution unit: the rack's managed power strip that feeds every device in it.
- KVM here means a keyboard-video-mouse switch, which lets one console control many servers.
- SAN, a storage area network: the shared storage pool described above.
Quick check
Which single line item dominates the on-premise CapEx in the case study?
With the hardware bought, the meter starts running. Every watt the servers draw is paid for, every watt of heat they make has to be pumped out by the cooling, and the room has a monthly rent. These recurring costs are the Operational expenditure (OpEx), and adding them to the Capital expenditure (CapEx) gives the Total cost of ownership (TCO).
Worked example
From watts to a three-year TCO
Server power
5 × 308 W = 1,540 W. Over a year, 1.54 kW × 8,760 h = 13,490.4 kWh. At 0.22 EUR/kWh that is 2,967.89 EUR/yr. The slide shows 2,962.
Cooling power
5 × 385 W = 1,925 W, which is 16,863 kWh a year, or 3,709.86 EUR/yr. The slide shows 3,702.
Rent
5 m² × 5 EUR/m² per month × 12 = 300 EUR/yr.
Annual OpEx
Using the slide's figures, 2,962 + 3,702 + 300 = 6,964 EUR/yr.
Three-year OpEx
8,885 + 11,106 + 900 = 20,891 EUR. The 8,885 comes from an unrounded annual power cost of about 2,961.7 EUR times three.
Total cost of ownership
68,824 + 20,891 = 89,715 EUR over three years.
Result
TCO 89,715 EUR: about 29,905 EUR a year, 2,492 EUR a month, or 498 EUR per server per month. The split is 76.7% CapEx and 23.3% OpEx.
Cooling costs more than computing
Look again at the two energy lines. Cooling, at 3,702 EUR/yr, costs more than the servers' own power at 2,962 EUR/yr. For every 308 W of IT load, the room spends another 385 W getting rid of the heat. The standard way to express this is power usage effectiveness, defined by The Green Grid as total facility energy divided by IT equipment energy. Here that ratio is at least (308 + 385) / 308 = 2.25, and that is before counting the UPS, lighting or any other overhead.
The Uptime Institute's surveys put the industry average PUE at about 2.50 in 2007 and about 1.56 in 2024, and large hyperscale sites run well below that average. A small server room cooled by a generic air conditioner is simply much less efficient than a purpose-built data center. That gap is one of the structural reasons the cloud can undercut on-premise running costs.
What the TCO leaves out
The 89,715 EUR is a clean, recomputable number, and that is its value. It is also a lower bound, because several real costs sit outside it. Read the list below against the cost buckets from the first concept.
| Cost | How the case study treats it | Why it matters |
|---|---|---|
| Personnel | Named on slide 29, never priced | Often the largest operating cost in traditional IT |
| Maintenance | Assumed equal on both sides | Cancels only if the assumption holds |
| Software licences | Assumed open source | Free licence, paid support and admin time |
| Cost of capital | Not modelled | Barroso et al. use rates of 7 to 12 percent |
| WAN bandwidth, off-site backup | Not modelled | A streaming firm pays for both every month |
| Refresh after year 3 | Outside the horizon | The cycle starts again |
| Idle capacity | Bought for peak | Armbrust et al. cite 5 to 20 percent average utilisation |
The energy lines are also sensitive to price. At 0.30 EUR/kWh instead of 0.22, three years of power plus cooling rise from about 20,033 EUR to about 27,318 EUR, about 36% more, exactly the price ratio 0.30 / 0.22, because energy cost is linear in price. When you build such a model for research, vary the inputs you are least sure about and report the range, not only the point estimate.
Quick check
Five servers draw 308 W each, all year, at 0.22 EUR/kWh. What is the annual energy cost?
Quick check
Why is the 89,715 EUR TCO best read as a lower bound?
Recall
Compute the yearly cooling energy cost of 5 cooling loads of 385 W each at 0.22 EUR/kWh.
5 × 385 = 1,925 W, times 8,760 h is about 16,863 kWh, times 0.22 is about 3,710 EUR (the slide shows 3,702), more than the servers' own power.
Recall
What is the 3-year TCO and its CapEx and OpEx split?
89,715 EUR. CapEx 68,824 EUR (about 77%) and OpEx 20,891 EUR (about 23%).
Recall
What implied PUE do the slide's cooling figures give, and how does it compare with the 2024 industry average?
(308 + 385) / 308 = 2.25 or more, against an industry average of about 1.56.
Recap
If you remember nothing else
- Legacy IT cost goes well beyond the purchase price. It covers hardware, housing, power and cooling, personnel and upgrades.
- CapEx is upfront asset spending depreciated over its life. OpEx is recurring running cost, and that includes maintenance.
- Case study: 5 dual-socket Xeon E5-2640 v2 servers, open-source software, a 3-year horizon, and maintenance assumed equal in both scenarios.
- CapEx is 68,824 EUR, and the 35,000 EUR SAN is half of it.
- OpEx is 6,964 EUR/yr (power 2,962, cooling 3,702, rent 300), or 20,891 EUR over 3 years.
- TCO = 68,824 + 20,891 = 89,715 EUR, about 498 EUR per server per month.
- Cooling of 385 W for every 308 W of IT load implies a PUE of at least 2.25, against an industry average of about 1.56.
- The figure is a lower bound, because personnel, licences, capital cost and bandwidth are excluded.
Sources
- The Datacenter as a Computer, 3rd edition (Barroso, Hölzle and Ranganathan)BookSpringerChapter 6, Modeling Costs: CapEx and OpEx, depreciation periods, cost of capital, and the people and licence costs of traditional IT.(opens in a new tab)
- IAS 16 Property, Plant and EquipmentDocsIFRS FoundationDefinition of PP&E and depreciation, and paragraph 12 on expensing day-to-day servicing.(opens in a new tab)
- Intel Xeon Processor E5-2640 v2 specificationsDocsIntel ARK8 cores, 16 threads, 2.00 GHz base, 20 MB cache, 95 W TDP, launched Q3 2013.(opens in a new tab)
- Above the Clouds: A Berkeley View of Cloud Computing (Armbrust et al.)PaperUC Berkeley EECS, UCB/EECS-2009-28CapEx to OpEx versus pay as you go, and 5 to 20 percent average server utilisation.(opens in a new tab)
- Global Data Center Survey 2024ArticleUptime InstituteAverage PUE of about 1.56 in 2024, down from about 2.50 in 2007.(opens in a new tab)
- PUE: A Comprehensive Examination of the Metric (White Paper 49)PaperThe Green Grid and ASHRAEPUE as total facility energy divided by IT equipment energy, with cooling and UPS in the facility total.(opens in a new tab)
- Cost Optimization Pillar: design principlesDocsAWS Well-Architected FrameworkAdopt a consumption model, and stop spending money on undifferentiated heavy lifting.(opens in a new tab)
Part 06: Cost case study, NIST definition and deployment models
Finishes the traditional IT versus AWS cost comparison, then introduces the NIST definition of cloud computing, its five essential characteristics and the four deployment models.
6 concepts, slides 35-44
Why this part matters
Part 05 priced owning the infrastructure. This part prices renting it, and then turns a cost table into a decision rule you can defend in an exam and in your research project's deployment choices.
The second half gives the vocabulary every later lecture assumes: the NIST definition of cloud computing, its five essential characteristics and its four deployment models. Exams ask for these lists word for word. More usefully, they are a test you can apply to any design, including an edge deployment, to decide whether it is really a cloud or just servers someone else racked.
By the end you can
- Read a cloud bill line by line, convert European number formats, and spot totals that do not match their rows.
- Compare on-premise and cloud TCO as cumulative curves, compute the savings percentage and the break-even year, and name the costs the slide model leaves out.
- Explain why the provider runs the hardware but the customer still owns availability and lock-in risk, and why on-premise IT survives.
- Quote the NIST SP 800-145 definition and apply its five essential characteristics as a test for “is this cloud?”.
- Classify a deployment as public, private, hybrid or community by who may use it, and explain why ownership and location do not decide private or community.
- Explain cloud bursting as the defining example of a hybrid cloud.
Take one virtual machine on AWS: an EC2 m2.xlarge instance with a 1 TB SSD block volume, running 24/7. The m2.xlarge is a previous-generation instance with 2 vCPU and 17.1 GiB of memory. Its bill has five lines, and none of them is paid before day one.
| Line item | Per month | One year | Three years |
|---|---|---|---|
| EC2 m2.xlarge + 1 TB SSD EBS | 145 | 1,740 | 5,398 |
| Data transfer | 75 | 900 | 2,707 |
| Load balancer | 22 | 264 | 780 |
| Object storage capacity | 28 | 336 | 991 |
| Object storage requests | 48 | 576 | 1,743 |
| Total as printed | 323 | 3,878 | 11,619 |
| Sum of the rows | 318 | 3,816 | 11,619 |
Compare that with slide 33. There, 68,824 EUR left the company before a single request was served. Here there is no start column at all. Every euro is Operational expenditure (OpEx), and every line is a resource the provider meters. That gives the general rule for any cloud bill:
This is the NIST characteristic called Measured service, seen from the side of the person paying. Because each line is metered separately, each can be scaled or switched off on its own. You can shrink the instance, move cold files to cheaper object storage, or put a cache in front of the requests, and only that line moves.
How AWS sets these prices
AWS's pricing page names three principles: pay as you go, save when you commit, and pay less by using more. The slide uses plain on-demand prices. AWS's pricing whitepaper (now marked as historical reference) says reservations save up to 75% and Spot capacity up to 90% against on-demand. So the slide's bill is the most expensive way to buy this workload, which matters when you compare it with owning servers. To price a design of your own, use the official AWS Pricing Calculator.
Worked example
Audit the bill
Convert the format
The slide uses European notation: the dot separates thousands. 5.398 € means 5,398 EUR and 11.619 € means 11,619 EUR, not eleven euros.
Sum the monthly rows
145 + 75 + 22 + 28 + 48 = 318. The printed total is 323, five euros higher.
Check the yearly column
Every yearly row is exactly 12 times its monthly row, and they add up to 3,816. The printed total is 3,878, which is close to 12 × 323 = 3,876. The total was most likely computed from the wrong monthly total.
Check the three-year column
Here the rows do add up to the printed 11,619, but they are not 36 times the monthly values: 36 × 145 = 5,220, not 5,398, and 36 × 22 = 792, not 780. The three-year figures came from a different price estimate.
Result
Use the printed totals when the lecture builds on them (slide 36 does), but do not expect to reproduce them from the rows. The honest single-VM yearly cost is somewhere between 3,816 and 3,878 EUR.
Recall
Why does slide 35's total not match its rows, and what is the right reading of 11.619 €?
The rows add up to 318 per month and 3,816 per year, but the totals say 323 and 3,878, so the printed totals contain an error. 11.619 € is European formatting for 11,619 euros.
Put the two options side by side. On-premise is a lump of 68,824 EUR at the start, then 6,964 EUR a year. AWS is nothing at the start and 19,364 EUR a year. After three years the slide reports 89,715 EUR against 58,093 EUR.
| Period | Traditional IT | AWS |
|---|---|---|
| Start | 68,824 | 0 |
| 1st year | 6,964 | 19,364 |
| 2nd year | 6,964 | 19,364 |
| 3rd year | 6,964 | 19,364 |
| Total after 3 years | 89,715 | 58,093 |
Where does 19,364 come from? Slide 35 priced one VM at about 3,878 EUR a year, and slide 33 sized the on-premise setup at five servers. 19,364 / 5 ≈ 3,873, so the AWS column is five VMs, one per server. Over three years, 5 × 11,619 = 58,095, within rounding of the printed total.
A verdict is a curve, not a number
The 35% depends entirely on stopping the clock at year three. The fair comparison writes both Total cost of ownership (TCO) figures as functions of time. On-premise starts high because of Capital expenditure (CapEx) and grows slowly. The cloud starts at zero and grows fast. Two lines like that must cross.
Worked example
Find the break-even year
Write both lines
On-premise: 68,824 + 6,964 t. AWS: 19,364 t.
Set them equal
68,824 = (19,364 - 6,964) t = 12,400 t, so t* ≈ 5.55 years.
Check on whole years
Years On-premise AWS Cheaper 3 89,716 58,092 AWS 5 103,644 96,820 AWS 5.55 ≈ 107,500 ≈ 107,500 Tie 6 110,608 116,184 On-premise Cumulative cost in EUR at selected horizons Result
AWS is cheaper in cumulative cost for the first five years. From year six on, on-premise wins, unless its hardware has to be replaced first. Servers are typically refreshed every three to five years, which resets the on-premise line with a new CapEx step before it ever reaches the crossing.
So the cloud wins when the horizon is short, when demand is uncertain or bursty, and when hardware would be refreshed before t*. Owning wins for steady load run for many years on hardware you keep.
Benefits the table does not price
Slide 37 lists what the money buys beyond the line items. The provider operates the hardware, keeps the infrastructure's Availability and upgrades it. Extra demand is met through Rapid elasticity. Armbrust and colleagues at Berkeley explain why that last point is worth so much. They call it cost associativity: “using 1000 EC2 machines for 1 hour costs the same as using 1 machine for 1000 hours”. They also note that “real world estimates of server utilization in datacenters range from 5% to 20%”, so owned capacity sits mostly idle. Most of all, the cloud removes the up-front commitment: you do not have to guess your peak before you know your users.
Why nobody deletes their own IT
- Some IT is always local. Laptops, office networks, printers and identity still need someone in the building.
- Some workloads must stay on premises. Tight latency, data residency rules or regulation can forbid sending data to a remote data center. This is the edge side of the E2C continuum from parts 01 and 02: edge resources sit close to the data source, often on the user's own premises.
- Lock-in is a real cost. Vendor lock-in grows with every proprietary service you adopt, and leaving later means paying outbound transfer and rewriting code.
Quick check
Using the slide 36 figures and ignoring hardware refresh, after roughly how long does on-premise become cheaper than AWS?
Recall
What is the three-year saving of AWS on slide 36, and why must you always quote it with its horizon?
(89,715 - 58,093) / 89,715 ≈ 35.2%. The two cumulative curves cross at t* ≈ 5.55 years, so the same comparison stopped at year six favours on-premise. The saving is a property of the horizon, not of the cloud.
Recall
What does “the provider is responsible for availability” leave to the customer?
Architecting for availability in the cloud, for example spreading across multiple availability zones. A single EC2 instance has only a 99.5% SLA, while the region-level SLA is 99.99%.
At 2 a.m. you open the AWS console on your phone, launch a VM, and it is running a minute later. An hour later you delete it and stop paying. Nobody at Amazon was woken up. Every part of that experience appears in one sentence that the whole field uses as its reference.
The NIST cloud definition, published by Mell and Grance as SP 800-145 in September 2011, reads:
Map the 2 a.m. VM onto it. A phone on any network is ubiquitous access. Launching without asking is on-demand. The VM came from a shared pool of hardware. It appeared in a minute and disappeared an hour later: rapidly provisioned and released. No one at the provider was involved: minimal interaction.
The 5-3-4 frame
The definition continues: the cloud model “is composed of five essential characteristics, three service models, and four deployment models”. Slide 38 draws them as three layers.
Service models are part 07. This part covers the five characteristics and the four deployment models. NIST later built its cloud reference architecture (SP 500-292) on the same definition, so the vocabulary carries through standards, procurement documents and papers.
Recall
Write the NIST definition of cloud computing from memory, then give the 5-3-4 frame.
“Cloud computing is a model for enabling ubiquitous, convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, servers, storage, applications, and services) that can be rapidly provisioned and released with minimal management effort or service provider interaction.” Then 5 essential characteristics, 3 service models (IaaS, PaaS, SaaS), 4 deployment models (public, private, community, hybrid).
Your department has a file server. To get space on it you email IT, wait two days, and receive a fixed quota that never shrinks. It runs on a VM, so it is “virtualized”. Is it a cloud? Test it against the five characteristics and it fails two at once: there is no self-service, and capacity does not grow or shrink with demand. It is virtualized hosting. Now test Google Drive or EC2. They pass all five.
The five characteristics are the test. Slide 39 gives one-line versions. The NIST text is more precise, and the precise words are where the exam marks are.
- On-demand self-service. A consumer can provision computing capabilities “unilaterally … as needed automatically without requiring human interaction with each service provider”.
- Broad network access. Capabilities are available over the network through “standard mechanisms” usable from “heterogeneous thin or thick client platforms”: phones, tablets, laptops, workstations.
- Resource pooling. The provider's resources serve many consumers in a “multi-tenant model”, “dynamically assigned and reassigned according to consumer demand”. The customer does not control the exact location but may specify it “at a higher level of abstraction (e.g., country, state, or datacenter)”. That is why you pick a region but never a rack.
- Rapid elasticity. Capabilities scale rapidly “outward and inward commensurate with demand”, and to the consumer they “often appear to be unlimited”.
- Measured service. Cloud systems “automatically control and optimize resource use by leveraging a metering capability” suited to the type of service (storage, processing, bandwidth, active user accounts). Usage “can be monitored, controlled, and reported, providing transparency for both the provider and consumer”. NIST adds in a footnote that metering is typically pay-per-use or charge-per-use. The bill in concept one is this characteristic made visible.
| Characteristic | NIST wording (condensed) | Everyday example | Without it |
|---|---|---|---|
| On-demand self-service | Provision capability unilaterally, as needed, without human interaction with each provider | Launching a VM from the console at 2 a.m. | Every request waits for a ticket and an administrator |
| Broad network access | Available over the network through standard mechanisms, from heterogeneous thin or thick clients | The same storage from a phone, laptop or script over HTTPS | Access only from one office network or one special client |
| Resource pooling | Multi-tenant pool, resources dynamically assigned and reassigned, location independent at a higher level | You choose a region, never a rack | Dedicated boxes per customer, idle most of the time |
| Rapid elasticity | Provisioned and released rapidly, outward and inward with demand, often appearing unlimited | An auto scaling group adds servers at peak and removes them after | You buy for the peak and pay for idle capacity all year |
| Measured service | Automatically controls and optimizes resource use through metering, and reports usage transparently to provider and consumer | The line items of the slide 35 bill | No pay-per-use, and no data to optimize cost with |
Quick check
A billing dashboard shows per-hour vCPU charges and per-GB transfer charges for every resource you use. Which essential characteristic does this show most directly?
Recall
Name the five NIST essential characteristics.
On-demand self-service, broad network access, resource pooling, rapid elasticity, measured service.
Three scenarios. Anyone with a credit card rents EC2 capacity in AWS's Bahrain region. A bank runs OpenStackfor its own business units only, either in its own data hall or in a hall it rents. And a student asks: if Saudi Aramco's internal cloud sits in a colocation facility, does it stop being private?
It does not, and the reason is the whole idea of a deployment model. Slide 40 summarizes it as “where the cloud infrastructure is deployed and who is allowed to use it”. Of those two, who may use it is what decides. NIST deliberately leaves ownership and, for most models, location open.
Public cloud
A Public cloud is provisioned “for open use by the general public”. It “may be owned, managed, and operated by a business, academic, or government organization, or some combination of them”, and it “exists on the premises of the cloud provider”, which means off premises for every user. This is the Hyperscaler world of AWS, Azure and Google Cloud. Armbrust et al. describe it as a cloud “made available in a pay-as-you-go manner to the general public”, and the provider is a Cloud service provider (CSP).
Private cloud
A Private cloud is provisioned “for exclusive use by a single organization comprising multiple consumers (e.g., business units)”. It may be owned, managed and operated by the organization, a third party, or both, and it “may exist on or off premises”. The bank's OpenStack deployment is private whether it runs in the bank's basement or in a rented hall, because only the bank's units can use it. The same goes for the colocated Aramco cloud.
| Model | Who may use it | Who may own and operate it | Where it is located |
|---|---|---|---|
| Public | The general public | A business, academic or government organization, or a mix | On the provider's premises |
| Private | One organization, with many internal consumers | The organization, a third party, or both | On or off premises |
| Community | A community of organizations with shared concerns | Members, a third party, or both | On or off premises |
| Hybrid | Whoever each component cloud serves | Each component keeps its own owner | Spans the locations of its components |
Quick check
A university runs OpenStack only for its own departments, on servers housed in a rented colocation facility. Which NIST deployment model is this?
Recall
Does a private cloud have to sit on the organization's premises?
No. NIST says it may be owned or operated by a third party and may exist on or off premises. Exclusive use by one organization is what defines it.
During Hajj season or Black Friday, an e-commerce site's traffic multiplies for a few days. The company runs its shop on its own private cloud, sized for a normal week. When that capacity hits its peak, the overflow is sent to EC2, and when the rush ends the rented capacity is released. That move has a name: Cloud bursting, and it is the textbook example of a hybrid cloud.
Hybrid cloud
NIST defines a Hybrid cloud as “a composition of two or more distinct cloud infrastructures (private, community, or public) that remain unique entities, but are bound together by standardized or proprietary technology that enables data and application portability (e.g., cloud bursting for load balancing between clouds)”. Two phrases carry the definition. The parts “remain unique entities”: the private cloud does not dissolve into the public one. And they are bound by technology that enables “portability”: the same application and data can run on either side.
AWS describes cloud bursting as “a configuration method that uses cloud computing resources whenever on-premises infrastructure reaches peak capacity”. It also sells the opposite direction: AWS Outposts brings “AWS infrastructure and services to virtually any on-premises or edge location for a truly consistent hybrid experience”, aimed at low latency, local data processing and data residency. Those are exactly the reasons the previous concept gave for keeping IT on premises, which is why hybrid designs show up all along the E2C continuum.
Community cloud
A Community cloud is provisioned “for exclusive use by a specific community of consumers from organizations that have shared concerns (e.g., mission, security requirements, policy, and compliance considerations)”. Like a private cloud, it may be owned and operated by community members, a third party, or both, and it may exist on or off premises.
AWS GovCloud (US) is a working example. It is offered to “verified U.S. government agencies and entities”, in “physically and logically isolated U.S. sovereign regions … operated by U.S. citizens on U.S. soil”. The agencies share it because they share the same compliance regime (FedRAMP High, ITAR), which is precisely NIST's “shared concerns”. NIST SP 800-146 discusses the trade-offs of each deployment model in more depth.
Quick check
Several hospitals jointly operate a cloud reserved for their patient-record workloads under the same health regulations. Which deployment model fits best?
Recall
What two properties make a composition of clouds “hybrid” under NIST?
The clouds remain unique entities, and they are bound by standardized or proprietary technology that enables data and application portability, for example cloud bursting.
Recap
If you remember nothing else
- A cloud bill is pure OpEx: unit price times metered quantity, line by line. The VM line (instance plus disk) was under half of the slide 35 bill.
- Slide 36: 89,715 EUR on-premise vs 58,093 EUR on AWS over three years, a saving of about 35.2%.
- TCO is a curve, not a number: break-even t* = CapEx / (cloud per year - on-premise OpEx per year) ≈ 5.55 years here.
- Unpriced cloud benefits are provider-run hardware, upgrades and elasticity. Availability above the infrastructure is still the customer's design job.
- Keep some traditional IT: latency, residency, regulation and vendor lock-in.
- NIST SP 800-145 has 5 essential characteristics, 3 service models and 4 deployment models.
- The five characteristics are on-demand self-service, broad network access, resource pooling, rapid elasticity and measured service.
- The deployment model is about who may use the infrastructure: anyone (public), one organization (private), organizations with shared concerns (community), or a composition (hybrid).
- Private and community clouds may sit on or off premises. A public cloud sits on the provider's premises.
- Hybrid clouds remain distinct units bound by portability technology. Cloud bursting is the canonical example.
Sources
- SP 800-145: The NIST Definition of Cloud ComputingDocsNIST (Mell and Grance, 2011)The definition, five characteristics, three service models and four deployment models. PDF: nvlpubs.nist.gov/nistpubs/Legacy/SP/nistspecialpublication800-145.pdf(opens in a new tab)
- SP 800-146: Cloud Computing Synopsis and RecommendationsDocsNIST (2012)Trade-offs of each deployment model.(opens in a new tab)
- SP 500-292: NIST Cloud Computing Reference ArchitectureDocsNIST (2011)(opens in a new tab)
- Above the Clouds: A Berkeley View of Cloud ComputingPaperArmbrust et al., UC Berkeley EECS-2009-28Cost associativity, 5 to 20 percent utilization, no up-front commitment, public vs private cloud.(opens in a new tab)
- How AWS Pricing Works: key principlesDocsAmazon Web ServicesMarked as historical reference. On-demand, reservations up to 75%, Spot up to 90%, outbound transfer charges.(opens in a new tab)
- AWS PricingDocsAmazon Web ServicesPay as you go, save when you commit, pay less by using more.(opens in a new tab)
- AWS Pricing CalculatorDocsAmazon Web Services(opens in a new tab)
- Amazon EC2 previous generation instancesDocsAmazon Web ServicesSpecification of m2.xlarge.(opens in a new tab)
- Shared Responsibility ModelDocsAmazon Web Services(opens in a new tab)
- Amazon Compute Service Level AgreementDocsAmazon Web Services(opens in a new tab)
- What is cloud bursting?DocsAmazon Web Services(opens in a new tab)
- AWS OutpostsDocsAmazon Web Services(opens in a new tab)
- AWS GovCloud (US)DocsAmazon Web Services(opens in a new tab)
Part 07: Service models: IaaS, PaaS and SaaS
Explains how IaaS, PaaS and SaaS split responsibility for the nine-layer stack between customer and provider.
4 concepts, slides 45-48
Why this part matters
Every system you design, and every platform you choose for a research prototype, starts with one decision: how much of the computing stack do you want to own? That choice sets how much operating work you take on, how much control you keep, how scaling happens, and who is accountable when something breaks. This part turns the three famous acronyms, IaaS, PaaS and SaaS, into one picture you can reason with. Exams ask you to classify a service and to say who manages a given layer, and this part gives you a reliable method for both.
By the end you can
- Explain a service model as the position of the provider and customer line on the nine-layer stack.
- Classify real services (EC2, Compute Engine, App Engine, App Service, Gmail, Microsoft 365) as IaaS, PaaS or SaaS.
- State who manages any given layer under IaaS, PaaS and SaaS, and quote the NIST wording for each model.
- Distinguish block, file and object storage, and the compute and networking resources IaaS rents.
- Argue the trade between control and automation, and explain why PaaS can autoscale.
- Identify the responsibilities the customer keeps in every model, and spot where the slides oversimplify them.
Suppose Majid's lab wants a small web app to track the papers everyone is reading. There are four ways to get it running. Option A: buy servers, rack them in a room on campus and run everything, the classic Legacy IT setup. Option B: rent virtual machines on Amazon EC2 and install the app on them. Option C: push the code to Google App Engine and let Google run it. Option D: subscribe to a ready-made paper tracker and just log in. The app the lab members see is the same in all four cases. What changes is how many layers of the stack the lab runs itself, and how many it hands to a Cloud service provider (CSP).
That observation is the whole idea of a service model. The lecture draws the stack as nine layers, from the bottom up: Data Center, Networking, Storage, Server, Virtualization, Operating Systems, Databases, Security and Applications. A service model is a statement about where one horizontal line sits on that stack. Everything below the line is the provider's job, and everything above it is yours. Reading the slide diagrams, the number of customer-managed layers takes four values, one per model:
| Layer | On-prem | IaaS | PaaS | SaaS |
|---|---|---|---|---|
| Applications | Customer | Customer | Customer | Provider |
| Security | Customer | Customer | Provider | Provider |
| Databases | Customer | Customer | Provider | Provider |
| Operating Systems | Customer | Customer | Provider | Provider |
| Virtualization | Customer | Provider | Provider | Provider |
| Server | Customer | Provider | Provider | Provider |
| Storage | Customer | Provider | Provider | Provider |
| Networking | Customer | Provider | Provider | Provider |
| Data Center | Customer | Provider | Provider | Provider |
The three model names come from the NIST cloud definition, NIST SP 800-145 (Mell and Grance, 2011), which defines exactly three service models: Software as a Service (SaaS), Platform as a Service (PaaS) and Infrastructure as a Service (IaaS). Each definition starts with the same phrase, "the capability provided to the consumer is", which tells you what NIST is really describing: what the customer is handed, and therefore what the customer is left to manage. The Berkeley "Above the Clouds" report makes the same point as a spectrum: utility computing offerings "will be distinguished based on the level of abstraction presented to the programmer and the level of management of the resources", with Amazon EC2 at one end and Google AppEngine at the other.
Moving the line up has a price and a payoff. The payoff is less operational work: fewer things to patch, back up, monitor and replace. The price is control. You can no longer pick the kernel, tune the database or install an unusual library, because those layers are no longer yours. Microsoft puts it in one sentence: "In an on-premises datacenter, you own the whole stack. As you move to the cloud, some responsibilities transfer to Microsoft."
- Applicationsyou
- Securityyou
- Databasesyou
- Operating Systemsyou
- Virtualizationprovider
- Serverprovider
- Storageprovider
- Networkingprovider
- Data Centerprovider
Classify a service
Amazon EC2 hands you a machine, so it is IaaS.
Layer split as the slides draw it. Application security and access stay partly yours under PaaS and SaaS; see the slide 47 and 48 errata.
In every model, SaaS included, the customer still owns your data, identities and accounts, access management (MFA, RBAC), client endpoints.
Recall
Name the nine stack layers bottom to top, and say where IaaS draws the provider and customer line.
Start with what launching an Amazon EC2 instance actually looks like. You pick an instance type, which AWS describes as one of "various configurations of CPU, memory, storage, networking capacity, and graphics hardware". You pick an Amazon Machine Image (AMI), a template that includes the operating system, say Ubuntu. You attach an EBS volume for persistent disk, and you set a security group, which AWS calls "a virtual firewall". A minute later you have what AWS calls "a virtual server", billed On-Demand per second with a minimum of 60 s. From that moment on, patching Ubuntu, installing PostgreSQL and hardening SSH are your jobs.
That is Infrastructure as a Service (IaaS) in practice. The provider runs the five bottom layers, Data Center, Networking, Storage, Server and Virtualization, and hands you a Virtual machine (VM). You run the four top layers: Operating Systems, Databases, Security and Applications. NIST phrases it as the capability to "provision processing, storage, networks, and other fundamental computing resources where the consumer is able to deploy and run arbitrary software, which can include operating systems and applications". The consumer "has control over operating systems, storage, and deployed applications; and possibly limited control of select networking components (e.g., host firewalls)." You get those resources through On-demand self-service and pay for them as a Measured service.
The resources an IaaS provider rents come in three families:
- Compute: CPU, GPU and RAM, enough to execute arbitrary tasks, from a web server to a model training job.
- Storage: block, file and object storage, three different shapes for data described below.
- Networking: virtual routers, switches and load balancers that connect your machines to each other and to the internet.
Three shapes of storage
Students often treat the three storage types as three speeds of the same disk. They are three different ways of addressing data. Block storage splits data into "fixed-sized blocks", each with its own address, and the server mounts the whole thing as a volume, exactly like a raw disk; AWS calls this EBS. File storage keeps data as files in hierarchical "directory trees" on network-attached storage that many clients can share at once; AWS calls this EFS. Object storage keeps data as "discrete units called objects", each with its own metadata, in "a flat namespace" where "the object's unique identifier provides the address". You reach objects through an HTTP API, not a mount; AWS calls this S3.
| Type | Unit | Addressing | Access | AWS example | Typical use |
|---|---|---|---|---|---|
| Block | Fixed-size block | Block address on a volume | Mounted as a raw disk by one server | EBS | OS boot disk, database files |
| File | File in a folder | Path in a directory tree | Shared over the network by many clients | EFS | Shared home directories, content shared by many VMs |
| Object | Object plus metadata | Unique ID in a flat namespace | HTTP API calls (PUT, GET) | S3 | Backups, media, datasets, static web assets |
What you still own
Worked example
Launching a database VM on EC2: who does what
Pick an instance type
You choose the balance of CPU, memory and network, for example a memory-heavy type for PostgreSQL. The provider guarantees the physical server behind it.Pick an AMI
You choose the operating system image. From now on the guest OS is yours, including every security patch.Attach an EBS volume
The provider runs the storage system and replicates the volume. You choose its size, format the file system and decide on backups.Configure the security group
The provider supplies the virtual firewall; you write its rules, for example allowing port 5432 only from the app servers.Install and patch
You install PostgreSQL, apply OS and database updates, manage users and deploy the application.Result
The customer runs four layers: Operating Systems, Databases, Security and Applications. The provider runs Data Center, Networking, Storage, Server and Virtualization.
The Berkeley report explains why this division is so natural for IaaS. An EC2 instance "looks much like physical hardware, and users can control nearly the entire software stack, from the kernel upwards". It exposes "raw CPU cycles, block-device storage, IP-level connectivity". That freedom has a cost: because the provider cannot know how your software behaves, it is "inherently difficult for Amazon to offer automatic scalability and failover". Scaling and recovery stay largely your design problem. Azure's equivalents are Azure Virtual Machines, Azure Disk Storage and virtual networks.
Recall
According to NIST SP 800-145, what does an IaaS consumer control?
Recall
Contrast block, file and object storage in one line each.
Quick check
A startup rents Amazon EC2 virtual machines and installs PostgreSQL on Ubuntu. Who is responsible for patching the Ubuntu operating system?
Quick check
An IaaS tenant needs storage that many virtual machines mount as one shared directory tree. Which storage type fits best?
Now deploy the lab's paper tracker, a small Python web app, to Google App Engine or Azure App Service. You upload the code and a short configuration file. The platform provisions servers, keeps the operating system and the Python runtime patched, routes traffic, and adds or removes instances as demand rises and falls. Azure adds continuous deployment from your repository and staging slots, so you can test a release before swapping it into production. That is the "development, testing and delivery" lifecycle the slide mentions, and you never logged into a server.
That is Platform as a Service (PaaS). The provider manages every layer up to and including the runtime, and you manage your application. NIST defines it as the capability to "deploy onto the cloud infrastructure consumer-created or acquired applications created using programming languages, libraries, services, and tools supported by the provider". The consumer does not control "network, servers, operating systems, or storage, but has control over the deployed applications and possibly configuration settings for the application-hosting environment." Google describes App Engine as "one of the fully managed, serverless platforms for developing and hosting web applications at scale" that scales instances automatically "based on demand". Microsoft describes App Service as a way to run "web applications, mobile back ends, and RESTful APIs without worrying about managing the underlying infrastructure", in .NET, Java, Node.js, Python or PHP, with "automatic OS/runtime management and patching".
Why PaaS can scale for you
On IaaS, the provider cannot safely clone your VM, because it does not know what state lives inside it. PaaS solves that by asking you to build the app in a particular shape. The Berkeley report notes that AppEngine enforces a "clean separation between a stateless computation tier and a stateful storage tier", and that its automatic scaling and high-availability mechanisms "rely on these constraints". Because any request can go to any copy of the stateless tier, the platform can add copies under load and remove them when traffic drops, which is Rapid elasticity handled by the provider. The price of that convenience is constraint: you work within the supported languages, libraries and storage services, and moving off the platform later can mean rewriting parts of the app, a form of Vendor lock-in.
| Task | IaaS (EC2) | PaaS (App Engine, App Service) |
|---|---|---|
| OS patching | Customer | Provider |
| Runtime upgrades (Python, Node.js) | Customer | Provider, within supported versions |
| Scaling | Customer designs it (or configures autoscaling) | Platform scales instances on demand |
| App code | Customer | Customer |
| Access control and app configuration | Customer | Customer, on provider tools (shared) |
| Freedom to install anything | Full, from the kernel up | Limited to what the platform supports |
Recall
Why can App Engine autoscale automatically, while EC2 leaves scaling and failover to you?
Quick check
A team deploys its web app on Azure App Service. An SQL injection flaw in the team's code leaks customer records. Who was responsible for preventing it?
KFUPM gives every student email through Microsoft 365. Students open Outlook in a browser or the phone app; administrators manage accounts and security settings in the Microsoft 365 admin center. Nobody at KFUPM installs, patches or backs up a mail server. That is Software as a Service (SaaS): the provider runs the entire application, and you use it.
NIST defines SaaS as the capability to "use the provider's applications running on a cloud infrastructure", accessible "from various client devices through either a thin client interface, such as a web browser (e.g., web-based email), or a program interface". That reliance on Broad network access is why SaaS works on any laptop or phone. The consumer does not manage the infrastructure, "or even individual application capabilities, with the possible exception of limited user-specific application configuration settings." AWS adds that SaaS is "in most cases" made of "third-party end-user applications": email, office suites, CRM, video calls.
What never leaves the customer
Here is the twist the slides miss. Even when the provider runs every layer of the stack, some responsibilities stay with the customer in every model, SaaS included. Microsoft's shared responsibility matrix makes this explicit: "For all cloud deployment types, you own your data and identities." Google says the same in its own words: "In SaaS, we own the bulk of the security responsibilities. You remain responsible for your access controls and the data that you choose to store in the application."
| Responsibility | On-prem | IaaS | PaaS | SaaS |
|---|---|---|---|---|
| Customer data | Customer | Customer | Customer | Customer |
| Configurations and settings | Customer | Customer | Customer | Customer |
| Identities and users | Customer | Customer | Customer | Customer |
| Client devices | Customer | Customer | Customer | Shared |
| Applications | Customer | Customer | Shared | Shared |
| Network controls | Customer | Customer | Shared | Microsoft |
| Operating system | Customer | Customer | Microsoft | Microsoft |
| Physical hosts, network, datacenter | Customer | Microsoft | Microsoft | Microsoft |
Read the matrix row by row. Customer data, configurations and settings, and identities and users stay "Customer" in all four columns. Microsoft adds a short list that is always retained by the customer, whatever the model: data, endpoints, accounts, and access management such as role-based access control, multifactor authentication and conditional access. Client devices show up as shared under SaaS because Microsoft can provide some device management capabilities, but endpoint protection and compliance on the laptop or phone are still yours. If a student account is phished because MFA was never enabled, the provider did not fail; the customer did.
Classifying any service
Exams like to name a product and ask for its model. One question settles almost every case: what am I handed, a machine, a runtime, or a finished app? A machine you install software on (EC2, Google Compute Engine, Azure Virtual Machines) is IaaS. A runtime you deploy code onto (App Engine, Azure App Service) is PaaS. A finished application you sign in to (Gmail, Microsoft 365) is SaaS.
Worked example
Classify and assign responsibility
Classify each service
EC2 hands you a virtual machine: IaaS. App Engine hands you a runtime for your code: PaaS. Gmail hands you a finished email application: SaaS.Who patches the operating system?
On EC2 the customer patches the guest OS. On App Engine and Gmail the provider does, because the OS sits below the line in both models.Who enables MFA for user accounts?
The customer in all three. Identities and access management stay with the customer in every model.Result
EC2: IaaS, customer patches, customer enables MFA. App Engine: PaaS, provider patches, customer enables MFA. Gmail: SaaS, provider patches, customer enables MFA.
Recall
Under SaaS, what responsibilities remain with the customer?
Quick check
A university uses Gmail through Google Workspace. A student account is phished because multifactor authentication was never enabled. Whose responsibility was that control?
Recap
If you remember nothing else
- A service model says where the provider and customer line sits on the stack. Customer-managed layers drop from 9 (on-prem) to 4 (IaaS), 1 (PaaS) and 0 (SaaS, per the slides).
- IaaS rents virtual hardware: compute (CPU, GPU, RAM), storage (block, file, object) and networking (routers, switches, load balancers). You run the OS and everything above it.
- PaaS runs your code on a managed runtime and handles OS patching and scaling. In exchange it constrains how the app is built.
- SaaS delivers a finished application through a browser or API. Inside the app the user adjusts only user-specific settings, but still owns the retained duties below.
- Moving up the models trades control for less operational work. EC2 and App Engine sit at the two ends of the spectrum.
- In every model the customer keeps its data, identities, access management and endpoints. Security is never fully outsourced.
- Classification shortcut: a machine means IaaS, a runtime means PaaS, a finished app means SaaS.
Sources
- The NIST Definition of Cloud Computing, SP 800-145 (Mell and Grance, 2011)DocsNISTExact definitions of SaaS, PaaS and IaaS.(opens in a new tab)
- Above the Clouds: A Berkeley View of Cloud Computing, EECS-2009-28 (Armbrust et al., 2009)PaperUC BerkeleyThe EC2 to AppEngine spectrum and why constraints enable autoscaling.(opens in a new tab)
- A View of Cloud Computing, CACM 53(4):50-58, 2010PaperACM(opens in a new tab)
- Cloud Programming Simplified: A Berkeley View on Serverless Computing, EECS-2019-3 (Jonas et al.)PaperUC Berkeley(opens in a new tab)
- Shared responsibility in the cloudDocsMicrosoft LearnThe responsibility matrix across on-prem, IaaS, PaaS and SaaS.(opens in a new tab)
- Shared Responsibility ModelDocsAWS(opens in a new tab)
- Shared responsibility and shared fateDocsGoogle Cloud(opens in a new tab)
- Student emailDocsKFUPMKFUPM student email runs on Microsoft 365 and Outlook.(opens in a new tab)
- Types of cloud computingDocsAWS(opens in a new tab)
- What is Amazon EC2?DocsAWS(opens in a new tab)
- Block vs file vs object storageDocsAWS(opens in a new tab)
- An overview of App EngineDocsGoogle Cloud(opens in a new tab)
- Overview of Azure App ServiceDocsMicrosoft Learn(opens in a new tab)
- Gmail (Google Workspace)DocsGoogle(opens in a new tab)
Part 08: Public cloud hyperscalers: AWS, GCP and Azure
Surveys the history and service portfolios of the three biggest hyperscalers.
5 concepts, slides 49-56
Why this part matters
A dissertation prototype or an industrial edge-cloud system almost always ends up on one of three platforms: Amazon Web Services, Google Cloud or Microsoft Azure. Knowing them is not about memorising product logos. It is about knowing what is truly comparable between them, because that decides how hard a migration is, what a design costs, and how locked in you become.
This part first asks why only three companies dominate a market with countless providers. It then follows each one from its origin to its present portfolio, and ends with a single map that translates any service between the three. Expect the exam to hand you a service on one provider and ask for its equivalent on another, and to name the block it belongs to.
By the end you can
- Explain why only a few companies can operate at hyperscale, using economies of scale and the Berkeley motives for becoming a provider.
- Trace how AWS, Google Cloud and Azure began, with key dates, and link each origin to its strengths.
- Read any provider portfolio through the six building blocks.
- Translate a service between the three providers with an equivalence table.
- Spot outdated or misspelled service names on the slides.
Picture two data centers in 2006. The first is a medium-sized facility with about 1,000 servers. The second is a very large one with about 50,000. The first pays about $95 per Mbit/s of network bandwidth per month. The second pays about $13. Storage costs the first $2.20 per GB per month and the second $0.40. One administrator in the first looks after about 140 servers, while in the second one person runs more than 1,000. These are James Hamilton's estimates, reported in the Berkeley “Above the Clouds” report, the same table part 01 used to explain why the cloud stays centralized. Here it explains who can afford to be the cloud.
That gap is the whole story of the Hyperscaler. The slides note that there are a great many public cloud service providers, yet they focus on three. The reason is economic. A provider that buys bandwidth, disks, power and staff at one fifth to one seventh of a medium firm's price can rent capacity by the hour below what that firm pays to own it, and still make a profit. This is Utility computing sold at a wholesale cost base. It is also why the Public cloud model rewards size so strongly: every new customer spreads the fixed cost of software and operations over more machines.
Worked example
What scale buys
Network
A service needs 100 Mbit/s of sustained bandwidth. The medium data center pays 100 × $95 = $9,500 per month. The very large one pays 100 × $13 = $1,300.Storage
10 TB stored for a month costs about 10,000 × $2.20 = $22,000 in the medium facility and 10,000 × $0.40 = $4,000 in the large one.People
Running 5,000 servers takes about 36 administrators at 140 servers each, but at most 5 at 1,000 servers each.Result
The raw prices give ratios of 95 / 13 ≈ 7.3 and 2.20 / 0.40 = 5.5. The report prints 7.1, 5.7 and 7.1, so quote the advantage as “about 5.7 to 7.1 times, as reported”. Other estimates in the same report put it at 3 to 5 times.
Who could become a provider, and why
Big data centers are necessary but not sufficient. Armbrust and colleagues argue that a provider also needs large-scale software infrastructure (MapReduce, the Google File System, Bigtable, Dynamo) and the operational expertise to run and defend it. In the early 2000s only a handful of internet companies had both. Given those assets, the report lists six motives; the four that matter for the hyperscalers are:
- Make a lot of money from economies of scale, turning a cost advantage into margin.
- Leverage existing investment, adding a revenue stream on top of data centers already built for internal use.
- Defend a franchise, giving existing enterprise customers a cloud path before a rival does.
- Attack an incumbent, establishing a beachhead before a single “800 pound gorilla” emerges (the report's example is Google App Engine).
The other two are leveraging customer relationships (as with IBM) and becoming a platform (as with Facebook). Keep the four above in mind. The next concepts show that each hyperscaler's origin story fits one of them especially well. The market today reflects the same concentration: in the second quarter of 2026 the three largest providers took almost two thirds of a fast-growing market.
Cloud infrastructure services market, Q2 2026 (Synergy Research Group)
- Quarter
- Q2 2026
- Total cloud infrastructure services revenue
- $143.4B, +43% year on year
- Amazon (AWS)
- 28%
- Microsoft (Azure)
- 20%
- Google (Google Cloud)
- 15%
- Top three combined
- 63%
Recall
Which four Berkeley motives for becoming a cloud provider matter most for the hyperscalers?
By the early 2000s, Amazon's engineering leaders noticed something expensive. Teams building shopping features were spending about 70% of their time on the same basic plumbing: storage, compute, databases. Amazon later called this undifferentiated heavy lifting: work every team must do that does not set the product apart. Leaders began to think about building “a shared layer of infrastructure services that all these teams can rely on”. They soon saw that outside developers had exactly the same problem, and at a 2003 offsite at Jeff Bezos's house the top managers decided to sell those services to everyone else.
This is the Berkeley motive “leverage existing investment” in its purest form. Werner Vogels, Amazon's CTO, said that many AWS technologies were first developed for Amazon's internal operations. The slide adds the design pressures behind that platform: absorbing huge traffic spikes such as the holiday season, supporting a global store, and moving services onto commodity Linux hardware and open-source software. Building for those pressures forced Amazon to make infrastructure into reusable services with clean interfaces, which is exactly what a customer of a Cloud service provider (CSP) needs.
The public launches came in quick succession. A message queue, SQS, appeared in beta in November 2004. In March 2006 Amazon announced object storage as “storage for the Internet”. In August 2006 EC2 opened in limited beta, renting a Virtual machine (VM) at 10 cents per virtual CPU hour. EC2 made the Infrastructure as a Service (IaaS) model concrete: start a server with an API call in minutes, add more when load rises (Rapid elasticity), and pay only for the hours used (Measured service).
From internal platform to public cloud
- 2003
- Offsite at Jeff Bezos's house: leaders decide the planned shared infrastructure services can be an external business
- November 2004
- Simple Queue Service (SQS) public beta. Fortune still counts S3 as AWS's first official launch
- 13 March 2006
- Amazon S3 announced as “storage for the Internet”
- 25 August 2006
- Amazon EC2 limited beta at 10 cents per virtual CPU hour
- Today
- Over 200 services, used in about 190 countries (AWS overview whitepaper)
Recall
Why did Amazon build the shared platform that became AWS?
A portfolio slide with twenty product icons looks like a list to memorise. It is easier to read it as a floor plan. Take a small web shop and place each piece of it on the AWS portfolio from slide 51.
Worked example
Placing a web shop on the six blocks
Networking
Route 53 resolves the shop's name. CloudFront, a Content delivery network (CDN), caches product images close to shoppers. Elastic Load Balancing spreads requests over the web servers, which sit inside a private VPC. Direct Connect would add a dedicated line from the company's own office or data center.Compute
The web servers run on EC2 instances, each a Virtual machine (VM). A service packaged as a Container could run on ECS instead.Database
Orders go to a managed database: DynamoDB for key-value access at scale, or RDS for a relational schema. Shopping sessions live in ElastiCache, an in-memory cache.Application Services
SQS queues new orders so payment and shipping can process them at their own pace. SNS sends the “your order has shipped” notifications. CloudSearch would power product search.Storage
Product photos sit in S3 as objects. The servers' disks are EBS block volumes. Invoices older than a year move to cheap archive storage (Glacier on the slide).Deployment and Management
CloudFormation describes the whole setup as code so it can be rebuilt. CloudWatch monitors it, and IAM controls who may touch what. Elastic Beanstalk could instead deploy the app as a Platform as a Service (PaaS) with much of this wiring done for you.Result
Every part of the shop lands in one of six blocks: Deployment and Management, Networking, Application Services, Compute, Storage, Database.
The general rule is that every hyperscaler portfolio fits the same six blocks. The slides draw Google Cloud and Azure with exactly the same layout. Learn the categories once and the vendor names become vocabulary: what changes between providers is the name on the box, not the job it does. That is the point of slide 52, which skips a detailed tour of every CSP because the big ones offer similar or comparable services. Building on Infrastructure as a Service (IaaS) or Platform as a Service (PaaS) on any of the three means picking a box in each block.
Comparable is not identical, though. Microsoft's own guide for AWS professionals says the two clouds offer similar products built independently, and warns that “not every matched service has exact feature-for-feature parity”. Even the containers that hold resources differ: AWS organises them under accounts, Azure under subscriptions. These differences in APIs, limits and behaviour are exactly where Vendor lock-in comes from.
Recall
Name the six portfolio blocks and one service per provider for Compute, Storage and Database.
In April 2008 a developer could upload a Python web application to Google App Engine and never see a server. Google ran it on its own infrastructure, scaled it as traffic grew, balanced the load and stored the data in Bigtable and the Google File System. The free preview gave each application 500 MB of storage, 200 million megacycles of CPU per day and 10 GB of bandwidth per day, enough for about 5 million page views a month. Raw virtual machines came much later: Compute Engine (often abbreviated GCE) was announced at Google I/O on 28 June 2012 and became generally available on 2 December 2013.
So Google entered the cloud from the opposite end to Amazon. AWS started with hardware-like building blocks and later added managed services. Google started with a fully managed Platform as a Service (PaaS) and only later offered Infrastructure as a Service (IaaS). The Berkeley report turns this into a general way to tell providers apart: by the level of abstraction they present to the programmer, the spectrum part 07 used to place IaaS and PaaS. EC2 sits at the low end, where an instance looks like physical hardware and you control nearly the whole software stack from the kernel up. App Engine sits at the high end, where Google enforced a clean split between a stateless request-reply compute tier and a stateful storage tier, and in return gave you automatic scaling and high availability. Azure, as launched, sat in between.
In Berkeley terms, App Engine is the report's example of “attack an incumbent”: its appeal lay in automating the scalability and load balancing features that developers would otherwise build themselves.
| Platform | What you control | What the provider automates | Constraint |
|---|---|---|---|
| Amazon EC2 | Nearly the whole software stack, from the kernel upwards | Virtualized hardware through a thin API of a few dozen calls | None on application type; scaling and failover are your problem |
| Microsoft Azure (2009) | Your code and language choice, not the OS or runtime | Network configuration, some failover and scaling, once you declare properties | Code compiled to the .NET Common Language Runtime |
| Google App Engine | Only the request handlers of a web application | Servers, load balancing, automatic scaling, Bigtable-backed storage | Stateless request-reply compute tier, rationed CPU per request |
Azure automated some scaling and failover only after you declared your application's properties, so the figure leaves scaling with you.
The constraint is the price of the convenience. App Engine could scale your application automatically precisely because it knew its shape. A general-purpose program that keeps state in memory or needs long computations did not fit. EC2 would run anything, but scaling and failover were left to you because the provider cannot know how your application replicates state. Neither end is better; the right point depends on the task, which is why every hyperscaler now offers the whole range.
Google's portfolio on slide 54 shows that convergence. App Engine is still there in Deployment and Management, now beside Compute Engine for raw VMs, a managed Kubernetes service for each Container, Cloud Storage, Persistent Disk, Bigtable and Cloud SQL. The Application Services block also lists Pub/Sub for messaging and Dataflow and Dataproc for data processing, a reminder of Google's strength in large-scale data. The six blocks are the same as Amazon's; only the names differ.
Quick check
On the 2009 Berkeley abstraction spectrum, how were the three platforms ordered from lowest to highest abstraction?
Picture a company that runs Windows Server in its own data center, signs its staff in through Active Directory, and works in Office every day. Its developers write .NET. For this company, Azure is the shortest road to the cloud: the identities, licences and skills it already has carry straight over. That is the Berkeley motive “defend a franchise”, and the report names Azure as its example: it “provides an immediate path for migrating existing customers of Microsoft enterprise applications”.
Ray Ozzie unveiled Windows Azure at the Professional Developers Conference on 27 October 2008. It became generally available on 1 February 2010 in 21 countries, with full service level agreements. On 25 March 2014 Microsoft announced it would rename the platform Microsoft Azure, and the new name took effect on 3 April 2014. The rename signalled that Azure was no longer only about Windows: by then it ran Linux virtual machines and open-source stacks too. The enterprise focus on slide 55 remains its strongest card, with deep integration across Windows, Active Directory and Microsoft 365.
One map for all three providers
Slide 56 completes the set, and the portfolio once again falls into the same six blocks. Putting the three slides side by side gives the single most useful table in this part. Read each row as one need, then look up the local name on each provider. Where the slide's label is outdated or misplaced, the table gives the current service and notes the slide's version.
| Block | Need | AWS (slide 51) | Google Cloud (slide 54) | Azure (slide 56) |
|---|---|---|---|---|
| Deployment and Management | PaaS app hosting | Elastic Beanstalk | App Engine | App Service (slide: Web Apps) |
| Deployment and Management | Infrastructure as code | CloudFormation | Deployment Manager (retired, use Infrastructure Manager) | Resource Manager |
| Deployment and Management | Monitoring | CloudWatch | Cloud Monitoring | Azure Monitor |
| Deployment and Management | Identity and access | IAM | Cloud IAM | Microsoft Entra ID with Azure RBAC (slide shows AD Domain Services) |
| Networking | DNS | Route 53 | Cloud DNS | Azure DNS |
| Networking | Private network | VPC | VPC (slide: Virtual Network) | Virtual Network |
| Networking | CDN | CloudFront | Cloud CDN | Azure Front Door (slide: Content Delivery Network) |
| Networking | Load balancing | Elastic Load Balancing | Cloud Load Balancing | Load Balancer, Traffic Manager |
| Networking | Dedicated link | Direct Connect | Cloud Interconnect | ExpressRoute (not on slide) |
| Application Services | Messaging | SQS, SNS | Pub/Sub | Queue Storage (Service Bus per Google's table) |
| Application Services | Search | CloudSearch (closed to new customers) | Not on slide | Azure AI Search |
| Compute | Virtual machines | EC2 | Compute Engine | Virtual Machines |
| Compute | Containers | ECS (slide); EKS for Kubernetes | GKE | AKS (slide: Containers) |
| Storage | Object | S3 | Cloud Storage | Blob Storage |
| Storage | Block | EBS | Persistent Disk | Managed Disks |
| Storage | Archive or file | S3 Glacier classes | Cloud Storage Archive class | File Storage (the slides mix archive and file) |
| Database | NoSQL | DynamoDB | Bigtable | Cosmos DB |
| Database | Relational | RDS | Cloud SQL | SQL Database |
| Database | In-memory cache | ElastiCache | Memorystore | Azure Managed Redis (slide: Redis Cache; Azure Cache for Redis is being retired) |
The equivalences are close enough to plan a design and to answer exam questions, but each row hides differences. Azure's Cosmos DB, for example, is a fully managed NoSQL and vector database that offers document, key-value, graph and table models, and a 99.999% Availability SLA for multi-region setups. That is broader than DynamoDB or Bigtable, so a migration between them is a redesign, not a copy. This is the practical face of Vendor lock-in on any Hyperscaler. Slide 56 also shows Communication Services, Azure's APIs for chat, SMS, voice and video, which has no box on the other two slides.
Recall
Put the hyperscaler launches in order with dates.
Recall
Why is “AD Domain Services” not Azure's answer to AWS IAM?
Recall
Why is moving from DynamoDB to Cosmos DB a redesign rather than a copy, even though both sit in the same block?
Quick check
A team runs Amazon DynamoDB and must move to Azure. Which box on the Azure portfolio is the closest counterpart?
Quick check
Which Berkeley motive best explains Microsoft launching Azure?
Quick check
Which label on the slides names a service that has since been renamed?
Recap
If you remember nothing else
- Hyperscale is economics: very large data centers buy network, storage and staff at roughly 1/5 to 1/7 of medium-sized prices.
- The Berkeley report gives six motives for becoming a provider; the hyperscalers fit leverage existing investment (Amazon), attack an incumbent (Google App Engine) and defend a franchise (Microsoft).
- The top three hold 63% of cloud infrastructure revenue (Q2 2026: Amazon 28%, Microsoft 20%, Google 15%).
- AWS grew out of Amazon's internal shared platform: S3 and EC2 in 2006.
- Google started at the PaaS end with App Engine (Apr 2008) and added Compute Engine (GA Dec 2013).
- Azure was announced Oct 2008, reached GA Feb 2010, was renamed Microsoft Azure in 2014, and plays to Microsoft's enterprise base.
- Providers differ by abstraction level: EC2 gives control (you manage from the kernel up), App Engine gives convenience (automatic scaling, constrained app shape), and Azure (2009) sat between.
- Every portfolio uses the same six blocks. Services are comparable, not identical, and the differences drive lock-in.
- Slides go stale: Container Engine is now GKE, Deployment Manager gives way to Infrastructure Manager, “Cosmo DB” is Cosmos DB, and IAM maps to Entra ID plus RBAC, not AD Domain Services.
Sources
- Above the Clouds: A Berkeley View of Cloud Computing (Armbrust et al.)PaperUC Berkeley EECS, UCB/EECS-2009-28Table 2 economies of scale, the six motives for becoming a provider, Vogels on AWS origins, and the EC2, Azure, App Engine abstraction spectrum.(opens in a new tab)
- SP 800-145: The NIST Definition of Cloud ComputingDocsNISTDefinitions of public cloud and of SaaS, PaaS and IaaS.(opens in a new tab)
- Q2 cloud market passes $143 billion, highest growth rate in eight yearsArticleSynergy Research GroupQ2 2026 market size, growth and provider shares.(opens in a new tab)
- The history of Amazon Web ServicesArticleFortuneThe 70 percent figure, the 2003 offsite and the shared infrastructure layer.(opens in a new tab)
- The AWS Blog: The First Five YearsDocsAWS News BlogAmazon Simple Queue Service introduced in November 2004, before S3 and EC2.(opens in a new tab)
- Announcing Amazon S3, Simple Storage ServiceDocsAmazon Web Services13 March 2006 announcement.(opens in a new tab)
- Amazon EC2 betaDocsAWS News BlogLimited beta on 25 August 2006 at 10 cents per virtual CPU hour.(opens in a new tab)
- Overview of Amazon Web Services: introductionDocsAmazon Web ServicesAWS began offering IT infrastructure services in 2006; over 200 services today.(opens in a new tab)
- Transition from Amazon CloudSearch to Amazon OpenSearch ServiceDocsAWS Big Data BlogCloudSearch closed to new customers on 25 July 2024.(opens in a new tab)
- Amazon Glacier developer guide: introductionDocsAmazon Web ServicesThe vault-based service no longer accepts new customers; use the S3 Glacier storage classes.(opens in a new tab)
- Introducing Google App EngineDocsGoogle App Engine BlogApril 2008 launch, features and preview quotas.(opens in a new tab)
- Google Compute Engine is now generally availableDocsGoogle Cloud BlogGeneral availability on 2 December 2013.(opens in a new tab)
- Introducing Certified Kubernetes and Google Kubernetes EngineDocsGoogle Cloud BlogContainer Engine renamed Google Kubernetes Engine in November 2017.(opens in a new tab)
- Deployment Manager deprecationDocsGoogle CloudSupport ended 1 April 2026; replacement is Infrastructure Manager.(opens in a new tab)
- Compare AWS and Azure services to Google CloudDocsGoogle CloudCross-provider service mapping used for the equivalence table.(opens in a new tab)
- Hitting send on the next 15 years of GmailArticleGoogleGmail launched on 1 April 2004.(opens in a new tab)
- Microsoft unveils Windows Azure at Professional Developers ConferenceArticleMicrosoftAnnouncement on 27 October 2008.(opens in a new tab)
- Windows Azure general availabilityArticleMicrosoftGeneral availability on 1 February 2010 in 21 countries.(opens in a new tab)
- Upcoming name change for Windows AzureDocsMicrosoft Azure BlogRename to Microsoft Azure announced 25 March 2014, effective 3 April 2014.(opens in a new tab)
- Azure for AWS professionalsDocsMicrosoft LearnIndependent implementations without exact parity, accounts versus subscriptions, and IAM mapped to Entra ID and RBAC.(opens in a new tab)
- New name for Azure Active DirectoryDocsMicrosoft LearnAzure AD renamed Microsoft Entra ID; Azure AD Domain Services renamed Microsoft Entra Domain Services.(opens in a new tab)
- Azure Cosmos DB overviewDocsMicrosoft LearnFully managed NoSQL and vector database, multiple data models, 99.999 percent multi-region SLA.(opens in a new tab)
- Retirement of Azure Cache for Redis: FAQDocsMicrosoft LearnAzure Cache for Redis tiers are being retired in favour of Azure Managed Redis.(opens in a new tab)
Part 09: Private cloud platforms: OpenStack and OpenNebula
Introduces self-hosted private IaaS platforms, focusing on the open-source OpenStack and OpenNebula projects.
4 concepts, slides 57-62
Why this part matters
Public hyperscalers are not the only way to get a cloud. Data-sovereignty rules, edge sites, telco 5G cores and university research clusters often need infrastructure as a service on hardware the organization owns. OpenStack and OpenNebula are the two open-source platforms this course names. Knowing how each is built lets you argue where it fits, sketch how a virtual machine is actually provisioned, and answer the exam questions about their origins and components.
By the end you can
- Explain what a private IaaS platform adds on top of plain virtualization, using the NIST essential characteristics.
- Recall the origins and stewards of OpenStack (2010, Rackspace and NASA, foundation in 2012, now OpenInfra) and OpenNebula (2005, Complutense, first release in 2008, OpenNebula Systems).
- Name the core OpenStack services and map each one to its AWS equivalent.
- Trace the API calls OpenStack makes to boot one virtual machine.
- Contrast OpenNebula's front-end and drivers design with OpenStack's many services, and spot the outdated details on slide 62.
Picture a research computing group at KFUPM that owns 40 servers. Today a student who needs a GPU virtual machine emails an administrator and waits about a week while someone carves out the machine by hand. Now install a private cloud platform on those same 40 servers. The student logs into a web dashboard, picks an image and a size, and the machine is running in about a minute. Nothing about the hardware changed. What changed is that the servers became a shared pool that users draw from themselves, and every draw is counted.
That is exactly the list of NIST traits from earlier in this lecture: On-demand self-service, Resource pooling and Measured service. NIST defines a Private cloud as cloud infrastructure provisioned for the exclusive use of a single organization, whether it sits on or off the organization's premises. A private IaaS platform is the software that makes this real. It turns a pile of servers, disks and switches into on-demand compute, networking and storage, exposed through an API and a dashboard. The research literature calls OpenNebula a "virtual infrastructure manager" for exactly this reason: it takes infrastructure you already have and turns it into a private or hybrid IaaS cloud.
The market for this software is large. Products span the whole stack from IaaS up to SaaS, and they split into two families. Some are open source, such as OpenStack andOpenNebula, which the rest of this part studies. Others are licensed commercially and charge fees per socket, per core or per node. Choosing open source does not make the cloud free: you still pay the capital expense for hardware and the staff to run it, as the TCO case study showed. What it buys is control, and freedom from Vendor lock-in at the platform layer.
| Question | Plain hypervisor | Private IaaS platform | Public IaaS |
|---|---|---|---|
| Who owns the hardware | You | You | The provider |
| Self-service API for users | No, an admin clicks | Yes | Yes |
| Multi-tenancy with quotas | No | Yes, per project | Yes, per account |
| Billing model | Bought up front | Bought up front, usage metered internally | Pay per use |
| Example | ESXi or KVM alone | OpenStack or OpenNebula | AWS EC2 |
Recall
Name three things a private IaaS platform adds on top of a plain hypervisor.
In 2010 two organizations each had half of a cloud. Rackspace, a hosting company, wanted its Cloud Files storage code to become an industry standard rather than a private product; that code became Swift. Anso Labs, contracting for NASA Ames, had published Nova, a Python "cloud computing fabric controller" for NASA's Nebula cloud. The two efforts merged. The first Design Summit met in Austin on 13 to 14 July 2010, the project was announced at OSCON on 21 July 2010, and the first release, named Austin, shipped on 21 October 2010.
The project's stated mission was a platform for both public and private clouds, and the official guide still calls OpenStack "an open source cloud computing platform for all types of clouds". What made it last is governance. In September 2012 the OpenStack Foundation took over stewardship, so no single vendor owns the code. In October 2020 it was renamed the Open Infrastructure Foundation (OpenInfra), and in 2025 it joined the Linux Foundation. The code is licensed under Apache 2.0 and ships on a fixed 6-month cycle with alphabetical names: 2025.1 Epoxy, 2025.2 Flamingo, 2026.1 Gazpacho, and 2026.2 Hibiscus, released on 30 September 2026.
OpenStack timeline
- July 2010
- Rackspace (storage) and NASA (compute) merge efforts; first Design Summit in Austin
- October 2010
- First release, Austin
- September 2012
- OpenStack Foundation formed as neutral steward
- October 2020
- Renamed Open Infrastructure Foundation (OpenInfra)
- 2025
- OpenInfra joins the Linux Foundation
- Every 6 months
- A new named release, for example 2025.1 Epoxy and 2026.1 Gazpacho
Recall
What did each founder contribute to OpenStack in 2010?
Quick check
Who stewards the OpenStack code today?
The fastest way to understand OpenStack is to follow one command. A user types the command below to start a small Ubuntu machine on a private network. That single line touches six separate services before a Virtual machine (VM) exists.
openstack server create --image ubuntu-24.04 --boot-from-volume 20 --flavor m1.small --network private demo-vmWorked example
Booting one VM
Authenticate with Keystone
The CLI sends the user's credentials to Keystone and gets back a token plus a catalog of service endpoints. Every later call carries that token.Call the Nova API
The CLI sends the create request to Nova, the compute service, which checks the token with Keystone.Ask Placement for a host
Nova asks Placement which compute hosts have enough free vCPU and RAM for the m1.small flavor, then its scheduler picks one.Look up the image in Glance
Nova asks Glance, the image service, for the ubuntu-24.04 image record: its format, size and where its bytes live. With --boot-from-volume, the host itself never downloads the image; Cinder copies it into the boot volume in the next steps.Get a port from Neutron
Nova asks Neutron for a port on the private virtual network, with an IP address and security rules.Attach a Cinder volume
Because of --boot-from-volume, Nova asks Cinder to create a 20 GB volume from the image and attaches it as the boot disk. Without that flag, Nova boots from an ephemeral disk and never calls Cinder.Start the guest
The hypervisor on the host, libvirt with KVM in most deployments, boots the machine. The application inside it can later store files in Swift.Result
Six services (Keystone, Nova, Placement, Glance, Neutron, Cinder) cooperate through their APIs to create one VM, plus the hypervisor that runs it.
The architecture rule behind the trace
The trace shows the design rule. OpenStack is a set of independent services, and each one owns one kind of resource and exposes it through a REST API. Every service authenticates callers through Keystone, and services talk to each other only through those public APIs. Inside a single service, its worker processes coordinate through an AMQP message broker such as RabbitMQ, and each service keeps its state in its own SQL database. Horizon, the web dashboard, has no special powers: like the CLI, the SDKs or a plain curl, it is simply another REST client.
This design is why OpenStack scales to large operators and why it is heavy to run. Each service can be scaled, upgraded and replaced on its own, but a production cloud means operating many daemons, a message broker and several databases. Its feature list mirrors a Hyperscaler almost one for one, which makes the AWS comparison the easiest way to remember it.
| OpenStack service | Resource it manages | AWS equivalent |
|---|---|---|
| Nova | Compute instances (VMs, bare metal) | EC2 |
| Swift | Objects over HTTP | S3 |
| Cinder | Block volumes | EBS |
| Neutron | Virtual networks, ports, routers | VPC |
| Keystone | Identity, tokens, service catalog | IAM |
| Glance | Boot images | AMI catalog |
| Horizon | Web dashboard | Management Console |
| Placement | Inventory and usage of hosts | No direct equivalent |
Recall
Which services does Nova require before it can boot an instance?
Recall
How do OpenStack services communicate with each other, and inside themselves?
Quick check
A tenant uploads an Ubuntu disk image that new instances will boot from. Which OpenStack service registers and serves it?
Quick check
A database needs a disk that survives VM deletion and attaches like a hard drive. Which service provides it?
Now take a mid-size company leaving VMware. It installs one OpenNebula front-end on a single VM, adds ten KVM hosts and a Ceph datastore, and gives its staff the Sunstone web interface. That is a working private cloud, with far fewer moving parts than the OpenStack service list in the previous concept.
The difference is architectural. OpenNebula is centralized. The front-end holds the whole control plane: the oned daemon, the scheduler and an XML-RPC API, with Sunstone and the CLI sitting on top as clients. Hosts are hypervisor nodes, usually running a type 1 hypervisor such as KVM, grouped into clusters. Everything else plugs in through drivers. Storage drivers cover NFS, Ceph, local disks and others; network drivers cover Linux bridges, 802.1Q VLANs, VXLAN and Open vSwitch. Slide 62's "customized cluster" is this picture: the same front-end can drive hosts on your own physical servers, in an on-premises data center, or on bare metal rented from a public cloud, which is how OpenNebula supports hybrid and edge deployments.
Slide 62's layers, read from the bottom
- Infrastructure
- Physical servers, on-premises data centers, or public-cloud bare metal
- Operating system
- A Linux distribution on every host (the slide lists CentOS, RHEL, Ubuntu, Fedora)
- Hypervisor
- Today KVM for VMs and LXC for system containers
- Networking
- Linux bridges, 802.1Q VLANs, VXLAN, Open vSwitch
- Storage
- Ceph, local LVM, NAS, StorPool, LINSTOR and others
Where it came from
OpenNebula is older than OpenStack. Ignacio M. Llorente and Rubén S. Montero started it in 2005 as a research project in the Distributed Systems Architecture group at Complutense University of Madrid, with the goal of managing virtual machines on distributed infrastructure. The first public release came in March 2008, under the Apache license. In 2010 the founders created a company, C12G Labs, to offer enterprise support; it was later renamed OpenNebula Systems.
| Aspect | OpenStack | OpenNebula |
|---|---|---|
| Origin | Rackspace and NASA | Complutense University of Madrid |
| First release | 2010 | 2008 |
| Steward | OpenInfra, under the Linux Foundation | OpenNebula Systems |
| Architecture | Many REST services plus AMQP and SQL | Single front-end plus drivers |
| Control-plane footprint | Large | Small |
| Main API | REST | XML-RPC |
| Hypervisors today | Mostly KVM | KVM and LXC |
| Typical user | Telcos and large operators | Enterprises leaving VMware, edge sites |
Recall
Where does OpenNebula's control plane live, and which API does it expose?
Quick check
Which virtualization drivers does a current OpenNebula 7.x release still support?
Quick check
A university team with three admins and twelve KVM hosts wants the fewest control-plane daemons. Which choice fits?
Recap
If you remember nothing else
- A private IaaS platform turns servers an organization owns into self-service, pooled, metered compute, network and storage for that one organization.
- OpenStack: started in 2010 by Rackspace (Swift) and NASA (Nova). Foundation since 2012, renamed OpenInfra in 2020, part of the Linux Foundation since 2025. Apache 2.0, a release every 6 months.
- Core OpenStack services and AWS equivalents: Nova (EC2), Swift (S3), Cinder (EBS), Neutron (VPC), Keystone (IAM), Glance (machine images), Horizon (console), Placement (no direct twin).
- OpenStack services talk to each other through REST APIs authenticated by Keystone, use an AMQP broker inside each service, and keep state in a SQL database.
- OpenNebula: a research project from 2005 at Complutense University of Madrid, first released in 2008. C12G Labs (2010) became OpenNebula Systems. Apache 2.0.
- OpenNebula runs one front-end (oned, XML-RPC API, Sunstone) that manages hypervisor clusters through storage and network drivers. Today it supports KVM and LXC.
- Slides 60 and 62 say "Public" in the title but belong to the private section.
Sources
- SP 800-145: The NIST Definition of Cloud ComputingBookNIST (Mell and Grance, 2011)(opens in a new tab)
- A Bit of OpenStack HistoryDocsOpenStack Project Team Guide(opens in a new tab)
- Logical architectureDocsOpenStack Install Guide(opens in a new tab)
- Get started with OpenStackDocsOpenStack Install Guide(opens in a new tab)
- Nova documentationDocsOpenStack(opens in a new tab)
- Swift documentationDocsOpenStack(opens in a new tab)
- Cinder documentationDocsOpenStack(opens in a new tab)
- Neutron documentationDocsOpenStack(opens in a new tab)
- Keystone documentationDocsOpenStack(opens in a new tab)
- Glance documentationDocsOpenStack(opens in a new tab)
- Placement documentationDocsOpenStack(opens in a new tab)
- Horizon documentationDocsOpenStack(opens in a new tab)
- OpenStack releasesDocsOpenStack(opens in a new tab)
- OpenStack software overviewDocsOpenInfra Foundation(opens in a new tab)
- The OpenStack mapDocsOpenInfra Foundation(opens in a new tab)
- The OpenStack Foundation becomes the Open Infrastructure FoundationArticleTechCrunch, 19 October 2020(opens in a new tab)
- OpenInfra joins the Linux FoundationArticleOpenInfra blog, 2025(opens in a new tab)
- OpenNebula overviewDocsOpenNebula 7.4 documentation(opens in a new tab)
- What's new in 6.10 (LXD and Firecracker drivers removed)DocsOpenNebula documentation(opens in a new tab)
- 7.0 compatibility guide (vCenter discontinued)DocsOpenNebula documentation(opens in a new tab)
- Networking overview (supported network drivers)DocsOpenNebula 7.0 documentation(opens in a new tab)
- OpenNebula: leading innovation in cloud computing managementArticleERCIM News 83(opens in a new tab)
- Virtual Infrastructure Management in Private and Hybrid CloudsPaperSotomayor, Montero, Llorente, Foster, IEEE Internet Computing, 2009(opens in a new tab)
- OpenNebula: A Cloud Management ToolPaperMilojicic, Llorente, Montero, IEEE Internet Computing, 2011(opens in a new tab)
Part 10: Service-level agreements and availability
Defines availability, maps SLA nines to allowed downtime, and shows how serial services combine into a compound SLA.
4 concepts, slides 63-68
Why this part matters
The earlier parts of this lecture asked where a service should run and who should own the hardware. This last part asks a sharper question: how often will that service actually be there when a user calls it? Every Cloud service provider (CSP) sells the answer as a single number with a row of nines, and every architecture you design chains several such numbers together.
Turning nines into minutes of downtime, and composing them correctly across a pipeline, is a calculation you should expect on the exam. You will also need it the moment your research system spans an edge site, a fog node and a cloud region: each hop is another dependency, and each dependency spends part of your reliability budget.
By the end you can
- Compute availability from uptime and downtime, from successful and valid requests, or from MTBF and MTTR.
- Convert any number of nines into allowed downtime per day, week, month and year, and state the period convention you used.
- Compute the compound SLA of services in series and contrast it with the availability of redundant replicas in parallel.
- Distinguish SLI, SLO and SLA, and read the commitments and service credits of a real provider SLA.
Take a campus course-registration API and watch it for a 30-day month, which is 720 h. During that month it is unreachable twice, for a total of 2 h. For the other 718 h it answered requests. The fraction of the month it could answer is 718 / 720 = 0.9972, or 99.72%. That fraction is the service's Availability.
The general rule only names the two pieces. Uptime is the time the application can serve requests, and Downtime is the time it cannot. Availability is the share of uptime in the total:
The denominator is simply the length of the observation period, . Rearranging gives the form you will use most often, because it turns a promised availability into an allowance of failure: . For the registration API, (1 − 0.9972) × 720 h ≈ 2 h, which is exactly where we started.
Counting requests instead of hours
Large providers rarely have a clean on or off switch. A service may be up for most users while one shard fails. So they often measure availability by requests rather than by time: successful requests divided by valid requests. Google's SRE book gives a concrete case. A system that serves 2.5 million requests a day with a daily target of 99.99% may fail up to 250 of them. AWS combines both views: it measures success in windows of 1 or 5 minutes and averages the windows into a monthly uptime percentage.
Estimating availability before you have measurements
At design time there is no uptime log yet. Instead you can estimate availability from two failure statistics: the mean time between failures (MTBF) and the mean time to recover (MTTR).
Worked example
A component that fails every 150 days
Put both statistics in one unit
MTBF is 150 d = 3600 h. MTTR is 1 h.Apply the estimate
A = 3600 / (3600 + 1) = 3600 / 3601 ≈ 0.99972.Result
About 99.97%. The formula shows the two levers you have: fail less often (raise MTBF) or recover faster (cut MTTR). Halving the recovery time to 30 min buys as much as doubling the time between failures.
Recall
A component has an MTBF of 30 days and an MTTR of 2 hours. What is its availability, and which two levers raise it?
Start with the most common promise in cloud marketing, 99.9%, often called “three nines”. Its unavailability is 1 − 0.999 = 0.001. Multiply that by a period and you get the Downtime it allows: 0.001 × 8760 h = 8.76 h a year, 0.001 × 720 h = 43.2 min in a 30-day month, and 0.001 × 168 h ≈ 10.1 min a week.
Now add one more nine. An Availability of 99.99% has an unavailability of 0.0001, ten times smaller, so every downtime figure shrinks tenfold: 52.56 min a year, 4.32 min a month. That is the whole rule. Google's SRE book calls each additional nine “an order of magnitude improvement”, and the cost of getting there typically rises much faster than tenfold.
The table below repeats the calculation for each common tier. Like the lecture slide, it assumes a 365-day year, a 30-day month and a 7-day week. Every cell on the slide recomputes correctly with those conventions, and the figures agree with the availability table in the Google SRE book. The last column gives the example workloads AWS associates with each tier, so you can feel what a nine buys.
| Availability | Per year | Per month | Per week | Per day | Typical workload (AWS) |
|---|---|---|---|---|---|
| 99% | 3.65 d | 7.2 h | 1.68 h | 14.4 min | Batch and ETL jobs |
| 99.9% | 8.76 h | 43.2 min | 10.1 min | 1.44 min | Internal tools |
| 99.95% | 4.38 h | 21.6 min | 5.04 min | 43.2 s | Online commerce |
| 99.99% | 52.56 min | 4.32 min | 1.01 min | 8.64 s | Video delivery |
| 99.999% | 5.26 min | 25.9 s | 6.05 s | 0.86 s | ATMs, telecom |
| 99.9999% | 31.5 s | 2.59 s | 0.605 s | 0.086 s | Not in the AWS table |
What makes a number an agreement
A Service-level agreement (SLA) is where availability becomes a business promise. The vocabulary has three layers. You measure an indicator, you set an objective for it, and you sign an agreement that attaches consequences to the objective. Google's SRE book draws the line sharply: if nothing happens when the target is missed, it is an SLO, not an SLA. Azure's guidance says the same in business terms: an SLA has financial and legal implications, while SLOs are internal targets.
SLI, SLO, SLA and the consequence that binds them
- SLI (indicator)
- The measurement itself, for example the fraction of valid requests that succeeded in a 5-minute window.
- SLO (objective)
- A target value for an SLI, such as 99.95% of requests succeed each month. It is internal: missing it triggers engineering work, not payments.
- SLA (agreement)
- A contract with the customer that names one or more SLOs and the consequences of missing them. Without a consequence, it is just an SLO.
- Service credit
- The usual consequence. Amazon EC2 commits to 99.99% monthly uptime per region and refunds 10% of the bill below that, 30% below 99.0% and 100% below 95.0%. Credits are the sole remedy.
Read the Amazon EC2 SLA as a worked case. It promises 99.99% monthly uptime for instances spread across a region, but only 99.5% for a single instance. The gap is deliberate: one machine can fail, and only an architecture that spreads across failure domains earns the higher number. If AWS misses the regional promise, you receive a credit of 10%, 30% or 100% of that month's bill depending on how far uptime fell. You never receive compensation for your own lost business.
Quick check
A service promises 99.9% availability. About how much downtime does it allow per 30-day month?
Quick check
Which statement best distinguishes an SLA from an SLO?
Recall
What turns an SLO into an SLA?
Recall
Convert 99.95% to allowed downtime per 30-day month.
A request rarely touches a single service. Picture a request that first passes through Service 1, a front end promising 99.99%, and then Service 2, a database promising 99.9%. The request succeeds only if both are up at that moment. If the two fail independently, the probability that both are up is the product of their availabilities: 0.9999 × 0.999 = 0.9989001, about 99.89%.
Enters the pipeline.
Must be up.
Must also be up.
Arrives only if both succeeded.
Notice where the answer lands. It is not between the two numbers. It is below the weaker one. That result is the Compound SLA, and it holds for any chain of hard dependencies, where the failure of one stage makes the whole request fail:
The inequality follows from the arithmetic. Every is at most 1, and multiplying by a number at most 1 can only keep a value the same or lower it. AWS puts the consequence plainly: a workload can be no more available than any of its hard dependencies.
A shortcut and a warning about scale
When every unavailability is small, the product is close to one minus their sum. Here the unavailabilities are 0.01% and 0.1%, which add to 0.11%, giving roughly 99.89%. The exact value is 99.89001%, so the shortcut is good enough for an exam sanity check.
The shortcut also shows why long chains hurt. Five services at 99.9% each give 0.999⁵ ≈ 0.99501, about 99.50%, which is roughly 43.7 h of downtime a year against 8.76 h for one service. AWS gives a gentler version: a 99.99% system that hard depends on two other 99.99% systems lands at about 99.97%. Every microservice, managed database or edge gateway you add in series spends part of the budget.
Redundancy runs the arithmetic the other way
The fix for a weak link is to stop needing it to be up. Place independent replicas behind automatic failover so that any one of them can serve. Now the request fails only if all replicas are down at once, so you multiply the unavailabilities instead:
Two 99.9% replicas give 1 − 0.001² = 0.999999, or 99.9999%. AWS notes a quick way to see this when every component is all nines: count the nines, 3 + 3 = 6. The same idea explains why the Amazon EC2 SLA commits to only 99.5% for a single instance but 99.99% for instances spread across a region: the regional promise assumes your workload can survive the loss of any one instance. The two numbers are business commitments, not outputs of the formula.
| Topology | Formula | Two at 99.9% | When it applies |
|---|---|---|---|
| Series (hard dependencies) | A = ∏ Aᵢ | 99.8001% | Every service must work for the request to succeed |
| Parallel (redundant replicas) | A = 1 − ∏ (1 − Aᵢ) | 99.9999% | Any one replica can serve, and failover is automatic |
Quick check
Service 1 offers 99.99% and Service 2 offers 99.9% in series. What is the compound SLA?
Quick check
Two independent replicas, each 99.9% available, sit behind failover so either can serve. What is the effective availability?
Recall
Why can a pipeline never be more available than its weakest hard dependency?
Recall
Estimate the compound availability of 99.95% and 99.9% in series without a calculator.
A compound availability of 99.89% is still an abstract percentage. To plan operations you need it in minutes of Downtime. Its unavailability is 0.11%, or 0.0011, and the rule from the previous concepts does the rest.
Worked example
The compound 99.89% as a downtime budget
Per day
0.0011 × 86,400 s ≈ 95.0 s, which is 1 min 35 s.Per week
0.0011 × 604,800 s ≈ 665 s, which is 11 min 5 s.Per 30-day month
0.0011 × 2,592,000 s ≈ 2851 s, which is 47 min 31 s.Per 365-day year
0.0011 × 31,536,000 s ≈ 34,690 s, which is about 9 h 38 min.Result
The two services together may be down for about 48 minutes a month and nearly 10 hours a year, while still meeting the compound figure.
Online calculators such as uptime.is, the tool shown on the slide, automate exactly this multiplication. They are convenient, but they hide the period convention. uptime.is uses an average Gregorian year and month (about 365.24 d and 30.44 d), so its monthly and yearly figures come out slightly larger than the hand calculation with 30 and 365 days. The table compares the three sources.
| Period | By hand (0.11%) | uptime.is/99.89 now | Slide 67 screenshot |
|---|---|---|---|
| Day | 95.0 s = 1 min 35 s | 1m 35s | 1m 35s |
| Week | 665 s = 11 min 5 s | 11m 5.3s | 11m 5.3s |
| Month | 47 min 31 s (30 days) | 48m 13s | 47m 49s |
| Quarter | 2 h 22 min 34 s (90 days) | 2h 24m 38s | 2h 23m 27s |
| Year | 9 h 38 min (365 days) | 9h 38m 33s | 9h 33m 48s |
From allowance to error budget
Operations teams turn this allowance into a working tool called an error budget. Google's SRE book defines it as the gap between the SLO and the measured performance: the budget of how much unreliability is remaining in the period. While budget remains, the team ships releases and runs experiments, each of which risks a little downtime. When the budget is spent, releases pause and effort shifts to reliability. The Service-level agreement (SLA) sets the outer limit; the error budget is how the team manages its way to it.
Azure's reliability guidance shows the same thinking from the design side. In its worked example the composite of a workload's components comes to about 99.45%, roughly 4 h of downtime a month. The business then signs a 99.90% SLA anyway, consciously accepting the risk of paying credits. That gap between what the architecture supports and what the contract promises is a business decision, and it is exactly what the compound arithmetic lets you see before you sign.
Recall
Three services at 99.95%, 99.9% and 99.99% run in series. Roughly how much downtime does the chain allow per 30-day month?
Recall
Why does uptime.is report 48m 13s per month at 99.89% when a hand calculation gives 47m 31s?
Recap
If you remember nothing else
- Availability = uptime / (uptime + downtime). Allowed downtime = (1 − A) × T for any period T.
- Each extra nine cuts the downtime budget tenfold: 99.9% allows 8.76 h per year and 43.2 min per 30-day month.
- An SLA is a contract with consequences (Amazon EC2 pays service credits). An SLO is an internal target, and an SLI is the measurement behind both.
- Hard dependencies in series multiply: 99.99% × 99.9% ≈ 99.89%, which is below even the weakest link.
- Independent redundant replicas combine as 1 − ∏(1 − Aᵢ): two 99.9% replicas give 99.9999%.
- Downtime calculators differ by period convention, so check them. The calculator screenshot on slide 67 is not even consistent with itself.
Sources
- Reliability Pillar: AvailabilityDocsAmazon Web Services, Well-Architected FrameworkDefinition, time and request based formulas, MTBF and MTTR estimate, hard dependencies, redundancy, tier examples.(opens in a new tab)
- Availability with dependenciesDocsAmazon Web Services, Availability and Beyond whitepaperProduct formula, rough-bound caveat, rules on reducing and choosing dependencies.(opens in a new tab)
- Amazon Compute Service Level AgreementDocsAmazon Web Services99.99% regional and 99.5% single-instance commitments, service credit tiers.(opens in a new tab)
- Architecture strategies for defining reliability targetsDocsMicrosoft Learn, Azure Well-Architected FrameworkSLA versus SLO, composite SLO as a product, the 99.45% composite and 99.90% SLA example.(opens in a new tab)
- Site Reliability Engineering, Chapter 3: Embracing RiskBookGoogle, O'Reilly 2016Time-based and request-based availability, nines as orders of magnitude, error budgets.(opens in a new tab)
- Site Reliability Engineering, Chapter 4: Service Level ObjectivesBookGoogle, O'Reilly 2016Definitions of SLI, SLO and SLA.(opens in a new tab)
- Site Reliability Engineering, Appendix A: Availability TableBookGoogle, O'Reilly 2016(opens in a new tab)
- NIST SP 500-307: Cloud Computing Service Metrics DescriptionDocsNIST, 2018Background on measurable cloud metrics used in SLAs.(opens in a new tab)
- SLA and uptime calculator: 99.89%Articleuptime.isThe calculator shown on slide 67.(opens in a new tab)