Majid Al-RaimiReference sheet

COE 558Lecture 02Reference

Reference sheet

E2C continuum and cloud computing compressed onto one page: the definitions, formulas and numbers to have in your head before a quiz or exam.

The E2C continuum and the fog

The continuum is a range of latencies: moving away from the client adds latency and adds capacity. The cloud stays centralized for economies of scale, so mini data centers (the fog) go out instead. Part 01: Latency continuum and the fog

RTTmin⁡=2dv,v≈2×108 m/sRTT_{\min} = \dfrac{2d}{v}, \qquad v \approx 2 \times 10^{8}\ \text{m/s}
Physics floor over fiber: 4,000 km gives at least 40 ms
LayerSpanServicesPriorityRTT floorExample
EdgeLANCompute, storageAutonomy~1 µs at 100 mObstacle detection
FogWANCompute, storage, networkingCoordination~0.5 ms at 50 kmTraffic signals, collision coordination
CloudInternetFull data centerGlobal connectivity~40 ms at 4,000 kmFleet-wide model training
Layer map (slide 9) with RTT floors
ResourceMediumVery largeRatio
Network$95 per Mbit/s per month$137.1
Storage$2.20 per GB per month$0.405.7
Administration~140 servers per admin>1,0007.1
Economies of scale, 2006 (Armbrust et al.): about 1,000 versus 50,000 servers
AspectFogEdge
StructureHierarchical, multi-layerA few peripheral devices
ServicesCompute, networking, storage, control, accelerationSpecific applications in a fixed logic location
LocationBetween end devices and the cloudThe end-device layer itself
Fog versus edge (NIST SP 500-325)

Roles of edge, fog and cloud

Edge serves the client, fog serves many edges, cloud serves everyone. Part 02: Edge, fog and cloud layers

LayerRoleScopeData horizonExample
EdgeImmediate processing near the clientOne client or siteMilliseconds (act and discard)Wake-word detection
FogIntermediate layer for a regionMany edges in a districtSeconds to daysTraffic management
CloudCentralized, large-scale processingEveryone, globallyMonths to yearsTraining an AI model
Roles of the three layers
  • Fog benefits: less edge-to-cloud traffic, geo-restricted data for compliance, shorter edge-to-edge paths.
  • Edge roles: WAN boundary, offload target, fast small-data processing, gateway that hides the fog (like a reverse proxy).
  • Cloud: big compute (HPC), big storage (big data), big networking (multi-cloud).
  • IoT tree: devices sense, gateways process locally, fog coordinates regionally, cloud analyses globally.
  • Fog filtering example: 200 × 4 Mbps = 800 Mbps raw becomes 3.2 Mbps of events, about 250× less.
d=v⋅td = v \cdot t
Distance travelled while waiting for a decision

Vehicle platoon

Setup
v = 30 m/s (108 km/h), gap about 10 m
Cloud, 60 ms
d = 30 × 0.06 = 1.8 m
RSU (fog), 5 ms
d = 30 × 0.005 = 0.15 m
Car or V2V, 1 ms
d = 30 × 0.001 = 0.03 m
Mapping
Cars are edge nodes, RSUs and the RSUC are fog, the cloud coordinates globally.
QuestionThis courseNIST SP 500-325
What edge namesLayer on the WAN boundaryEnd devices and their users
Where a gateway sitsEdge layerAmong the fog nodes
Use it forExam answers in this coursePapers citing NIST SP 500-325
Where the edge boundary sits

Legacy IT and virtualization

A VM is an efficient, isolated duplicate of a real machine. Part 03: Legacy IT and virtualization basics

Legacy IT stack

Bottom to top
Data center, networking, storage, server, virtualization, operating system, databases, security, applications
Legacy IT
The customer manages all nine layers.
McCarthy (1961)
Computing sold like a utility: shared capacity, pay per use, network access.

Popek and Goldberg's three properties

Equivalence
Programs see an environment essentially identical to the real machine.
Efficiency
A statistically dominant subset of instructions runs directly on the real CPU. Rules out emulators and simulators.
Resource control
The VMM keeps complete control of real resources; a guest cannot reach what it was not given.
Sensitive⊆Privileged\text{Sensitive} \subseteq \text{Privileged}
Theorem 1: a sufficient condition for a trap-based VMM
AspectVirtualizationEmulationSimulation
Core ideaLogical copies of physical resourcesMimics a hardware/software interfaceModels a system's behaviour
Runs nativelyMost guest instructionsNothing, all translatedThe model only
Same ISAYesNoNot applicable
SpeedClose to nativeMuch slowerWhatever the model needs
ExampleVMware, VirtualBoxBlueStacks, console emulatorFlight simulator
Virtualization, emulation and simulation
Full (unmodified guest)Para (modified guest)
Type 1, bare metalVMware ESXi, Xen HVMXen PV
Type 2, hostedVirtualBox, VMware WorkstationParavirtual drivers in hosted products
Two independent axes: placement (rows) by guest modification (columns)

Hypervisors and containers

Cloud providers use Type 1: less overhead is more sellable capacity, and a thinner layer between tenants is stronger isolation. Part 04: Hypervisors and containers

PropertyType 1 (bare metal)Type 2 (hosted)
Runs onThe hardware, as the lowest layerA process of a host OS
Overhead (slide)≈5%≈10%
Resource managementDirectIndirect, through the host OS
AdministrationNeeds more skillEasy, like a desktop app
IsolationStrongerWeaker: host OS compromise exposes guests
ExamplesESXi, Hyper-V, KVM (debated), AWS NitroVirtualBox, VMware Workstation, Parallels
Type 1 versus Type 2
PropertyVMContainer
KernelOne guest kernel per VMShared host kernel
Start-upSeconds to minutesMilliseconds to about a second
Isolation boundaryHypervisor and virtual hardwareSystem call interface of a shared kernel
OS familyAny the virtual hardware supportsSame as the host kernel only
Image sizeGigabytesMegabytes
Virtual machine versus container
AspectNamespacesCgroups
QuestionWhat can a process see?How much can it use?
MechanismOwn view of a kernel resourceQuotas, weights, accounting, freezing
Types or functionsPID, network, mount, UTS, IPC, user, cgroup, timeLimiting, prioritization, accounting, control
Example files or flagsCLONE_NEWPID, CLONE_NEWNETcpu.max, cpu.weight, memory.max, cgroup.freeze
Namespaces versus cgroups
CPUs=quotaperiod=150000100000=1.5si=wi∑jwj\begin{gathered} \text{CPUs} = \frac{\text{quota}}{\text{period}} = \frac{150000}{100000} = 1.5 \\ s_i = \frac{w_i}{\sum_j w_j} \end{gathered}
cpu.max is a hard ceiling; cpu.weight is a share that matters only under contention

On-premise cost and TCO

Five dual-socket Xeon E5-2640 v2 servers, open-source software, a 3-year horizon. Part 05: On-premise TCO case study

AspectCapExOpEx
When paidUpfront and lumpyRecurring and smooth
AccountingCapitalised, then depreciatedExpensed in the period
ExamplesServers, SAN, switches, racksPower, cooling, rent, salaries, maintenance
Cloud analogueNone, the provider owns the hardwarePay as you go
CapEx versus OpEx
CapEx=∑iqi pi\text{CapEx} = \sum_i q_i \, p_i
Quantity times unit price over every asset bought at t = 0
Eyr=N⋅P⋅87601000 kWhcostyr=Eyr×price\begin{gathered} E_{\text{yr}} = \frac{N \cdot P \cdot 8760}{1000}\ \text{kWh} \\ \text{cost}_{\text{yr}} = E_{\text{yr}} \times \text{price} \end{gathered}
Yearly energy and its cost
TCO3 yr=CapEx+3×OpExyr\text{TCO}_{3\,\text{yr}} = \text{CapEx} + 3 \times \text{OpEx}_{\text{yr}}
Three-year total cost of ownership

Case study numbers (slides 33 and 34)

Servers
5 × 3,500 = 17,500 EUR
SAN
35,000 EUR (≈50.9% of CapEx, 39.0% of TCO)
Switches
4 × 3,677.50 = 14,710 EUR
Facilities, cooling kit
897 + 717 EUR
CapEx
68,824 EUR
Power
5 × 308 W, 13,490 kWh at 0.22 EUR/kWh: 2,962 EUR/yr (slide)
Cooling
5 × 385 W: 3,702 EUR/yr (slide)
Rent
5 m² × 5 EUR × 12 = 300 EUR/yr
OpEx
6,964 EUR/yr, 20,891 EUR over three years
TCO, 3 years
89,715 EUR, about 498 EUR per server per month (76.7% CapEx)
PUE=EfacilityEIT≥308+385308=2.25\text{PUE} = \frac{E_{\text{facility}}}{E_{\text{IT}}} \geq \frac{308 + 385}{308} = 2.25
Implied PUE, against an industry average near 1.56 (2024)

Cloud cost and the break-even

A cloud bill is pure OpEx: unit price times metered quantity. Part 06: Cost case study and NIST models

Cmonth=∑rpr⋅qrCyear=12 Cmonth\begin{gathered} C_{\text{month}} = \sum_{r} p_r \cdot q_r \\ C_{\text{year}} = 12\, C_{\text{month}} \end{gathered}
Cloud cost per resource line
PeriodTraditional ITAWS
Start68,8240
Each year6,96419,364 (5 VMs)
After 3 years89,71558,093
After 5 years103,64496,820
After 6 years110,608116,184
Traditional IT versus AWS, cumulative EUR
savings=89715−5809389715≈35.2%\text{savings} = \frac{89715-58093}{89715} \approx 35.2\%
Three-year saving of AWS
t∗=CapExCcloud−OpExon=6882419364−6964≈5.55 years\begin{aligned} t^{*} &= \frac{\text{CapEx}}{C_{\text{cloud}}-\text{OpEx}_{\text{on}}} \\ &= \frac{68824}{19364-6964} \approx 5.55\ \text{years} \end{aligned}
Break-even year; servers are usually refreshed before it

NIST SP 800-145: 5-3-4

Five essential characteristics, three service models, four deployment models. Part 06: Cost case study and NIST models

CharacteristicMeaningExample
On-demand self-serviceProvision unilaterally, no human at the providerLaunch a VM from the console at 2 a.m.
Broad network accessStandard mechanisms, heterogeneous clientsSame storage from phone, laptop or script
Resource poolingMulti-tenant, location independentYou pick a region, never a rack
Rapid elasticityScale out and in with demand, appears unlimitedAuto scaling group at peak
Measured serviceMetered, reported to provider and consumerLine items of the slide 35 bill
Five essential characteristics
ModelUsersLocation
PublicThe general publicOn the provider's premises
PrivateOne organizationOn or off premises
CommunityOrganizations with shared concernsOn or off premises
HybridComposition of distinct clouds bound by portability technologySpans its components
Four deployment models: who may use it
  • Cloud bursting is the canonical hybrid example.
  • Keep some traditional IT for latency, data residency, regulation and vendor lock-in.

Service models: IaaS, PaaS, SaaS

Machine means IaaS, runtime means PaaS, finished app means SaaS. Part 07: Service models

customer-managed layers9→4→1→0\begin{gathered} \text{customer-managed layers} \\ 9 \rightarrow 4 \rightarrow 1 \rightarrow 0 \end{gathered}
On-prem, IaaS, PaaS, SaaS as the slides draw it
LayerOn-premIaaSPaaSSaaS
ApplicationsCustomerCustomerCustomerProvider
SecurityCustomerCustomerProviderProvider
DatabasesCustomerCustomerProviderProvider
Operating systemsCustomerCustomerProviderProvider
VirtualizationCustomerProviderProviderProvider
Server, storage, networking, data centerCustomerProviderProviderProvider
Who manages each layer (slide diagrams)
TypeAddressingAccessAWSTypical use
BlockFixed-size block on a volumeMounted as a raw disk by one serverEBSBoot disk, database files
FilePath in a directory treeShared over the networkEFSShared home directories
ObjectUnique ID in a flat namespaceHTTP API (PUT, GET)S3Media, backups, data lakes
Block, file and object storage

Public cloud hyperscalers

Hyperscale is economics: roughly 1/5 to 1/7 of medium-sized unit prices. Part 08: Public cloud hyperscalers

Market and history

Market, Q2 2026
$143.4B, +43% year on year
Shares
Amazon 28%, Microsoft 20%, Google 15%, top three 63%
AWS
S3 (March 2006), EC2 (August 2006), from Amazon's internal platform
Google Cloud
App Engine PaaS (April 2008), Compute Engine GA (December 2013)
Azure
Announced October 2008, GA February 2010, renamed Microsoft Azure in 2014
Spectrum
EC2 gives control (kernel up), App Engine gives convenience (constrained app shape), Azure 2009 sat between.
BlockNeedAWSGoogle CloudAzure
ComputeVirtual machinesEC2Compute EngineVirtual Machines
ComputeKubernetesEKS (slide: ECS)GKEAKS
StorageObjectS3Cloud StorageBlob Storage
StorageBlockEBSPersistent DiskManaged Disks
DatabaseNoSQLDynamoDBBigtableCosmos DB
DatabaseRelationalRDSCloud SQLSQL Database
NetworkingPrivate networkVPCVPCVirtual Network
NetworkingDNSRoute 53Cloud DNSAzure DNS
NetworkingCDNCloudFrontCloud CDNAzure Front Door
Deployment and ManagementInfrastructure as codeCloudFormationInfrastructure ManagerResource Manager
Deployment and ManagementIdentity and accessIAMCloud IAMEntra ID with Azure RBAC
Application ServicesMessagingSQS, SNSPub/SubQueue Storage, Service Bus
Service equivalence by block (comparable, not identical)

Private cloud platforms

A private IaaS platform turns owned servers into self-service, pooled, metered resources for one organization. Part 09: Private cloud platforms

ServiceManagesAWS equivalent
NovaCompute instancesEC2
SwiftObjects over HTTPS3
CinderBlock volumesEBS
NeutronVirtual networksVPC
KeystoneIdentity, tokens, catalogIAM
GlanceBoot imagesAMI catalog
HorizonWeb dashboardManagement Console
PlacementHost inventory and usageNo direct equivalent
Core OpenStack services
AspectOpenStackOpenNebula
OriginRackspace and NASAComplutense University of Madrid
First release20102008
StewardOpenInfra, under the Linux FoundationOpenNebula Systems
ArchitectureMany REST services, AMQP, SQLSingle front-end (oned) plus drivers
Main APIREST, authenticated by KeystoneXML-RPC
Hypervisors todayMostly KVMKVM and LXC
LicenseApache 2.0Apache 2.0
OpenStack versus OpenNebula

SLAs and availability

Each extra nine cuts the downtime budget tenfold. Part 10: SLAs and availability

A=UptimeUptime+DowntimeD=(1−A)×T\begin{gathered} A=\frac{\text{Uptime}}{\text{Uptime}+\text{Downtime}} \\ D=(1-A)\times T \end{gathered}
Availability and allowed downtime over period T
Aest=MTBFMTBF+MTTRA_{est}=\frac{\text{MTBF}}{\text{MTBF}+\text{MTTR}}
Estimate from failure frequency (MTBF) and mean time to recover (MTTR)
AvailabilityPer yearPer monthPer day
99%3.65 d7.2 h14.4 min
99.9%8.76 h43.2 min1.44 min
99.95%4.38 h21.6 min43.2 s
99.99%52.56 min4.32 min8.64 s
99.999%5.26 min25.9 s0.86 s
Allowed downtime (365-day year, 30-day month)

SLI, SLO, SLA

SLI
The measurement, for example the share of valid requests that succeeded.
SLO
An internal target for an SLI. Missing it triggers engineering work, not payments.
SLA
A contract naming SLOs and the consequence of missing them.
EC2 example
99.99% monthly per region (99.5% single instance). Credits 10%, 30% below 99.0%, 100% below 95.0%. Credits are the sole remedy.
Aseries=∏i=1nAi≤min⁡iAiAparallel=1−∏i=1n(1−Ai)\begin{gathered} A_{\text{series}}=\prod_{i=1}^{n} A_i \le \min_i A_i \\ A_{\text{parallel}}=1-\prod_{i=1}^{n}(1-A_i) \end{gathered}
Series multiplies; independent replicas fail only together
TopologyFormulaResult
Series (hard dependencies)A = ∏ Aᵢ ≤ min Aᵢ99.8001%
Parallel (redundant replicas)A = 1 − ∏ (1 − Aᵢ)99.9999%
Two components at 99.9%

Slide errata

Answer with the corrected fact, and mention the slide's version when a question depends on it.

What the slides get wrong

Slide 10
"Mini Data Cener" should read Mini Data Center. Part 02
Slide 20
Full and para (guest modified?) are a separate axis from Type 1 and Type 2 (where it runs). VM/370 is 1972, not 1970/71. Part 03
Slide 22
vSphere is the suite; the Type 1 hypervisor is ESXi. KVM's label is debated. Part 04
Slide 24
Type 2 has less isolation than Type 1. 5% and 10% are workload-dependent approximations. Part 04
Slide 25
Virtual hardware is shared only when the container host is a VM. The SCI is the kernel entry point; the kernel manages hardware. Part 04
Slides 26, 27
Linux has eight namespace types. One namespace is not a container: a container is several namespaces plus cgroups plus a root filesystem. Part 04
Slide 32
Routine maintenance is OpEx (IAS 16 paragraph 12), not CapEx. Part 05
Slides 33, 34
European number format (17.500 € is 17,500 EUR). KVM is a keyboard-video-mouse switch. Rent is per m² per month. Part 05
Slide 35
Monthly rows sum to 318, not 323; yearly rows to 3,816, not 3,878. Part 06
Slide 36
The AWS column is five VMs (19,364 ≈ 5 × 3,873), never stated. Part 06
Slides 42, 44
"Comany" should read Company. Linking private clouds to the community cloud is a hybrid composition. Part 06
Slides 47, 48
The customer always keeps data, identities, access management and endpoints, even in PaaS and SaaS. Part 07
Slides 53, 54, 56
Gmail (2004) predates App Engine. Container Engine is GKE, Deployment Manager gives way to Infrastructure Manager, "Cosmo DB" is Cosmos DB, IAM maps to Entra ID plus RBAC. Part 08
Slides 60, 62
Titled "Public" but OpenStack and OpenNebula are private platforms. OpenNebula dropped Firecracker, LXD and vCenter drivers. Part 09
Slide 66
The second box should read Service 2, at 99.9%. Part 10
Slide 67
The calculator screenshot is internally inconsistent; 0.11% of a 365-day year is about 9 h 38 min. Part 10