COE 558Lecture 02Part 03
From legacy IT to virtualization
Starts the cloud computing section with the utility computing vision and the legacy IT stack, then defines virtualization, its history and its main forms.
- Concepts
- 6
- Slides
- 15-21
- Reading
- 36 min
Why this part matters
This part opens the cloud half of the lecture. Every cloud product you will price in parts 05 and 06, classify in part 07 or deploy in parts 08 and 09 is, at heart, a way of handing some layers of the classic IT stack to someone else. Before you can say which layers move, you need to know what the stack is and which technology lets one provider sell slices of a machine to many customers at once.
That technology is virtualization, and it is what turned a 1961 economic idea into a business. For your research and for the edge systems you will design, the trade-offs between a virtual machine, an emulator and a container decide how much isolation you get, how much overhead you pay and how portable your workload is. This part builds those distinctions carefully, and it flags two places where the slides blur them.
By the end you can
- Explain McCarthy's utility computing vision and map it to modern cloud characteristics.
- List the layers of legacy IT and say who manages each one.
- Define a virtual machine and explain transparency using Popek and Goldberg's three properties.
- Distinguish virtualization, emulation and simulation with an example of each.
- Classify virtualization techniques on two independent axes: guest modification and hypervisor placement.
- Explain which resources are isolated and which are shared in the hardware virtualization stack.
Look at your electricity bill. You never bought a power station, you never hired the engineers who run it, and you can plug in a kettle or a server without asking anyone. You pay for the kilowatt-hours you actually used, and the grid is reached through a socket in the wall. Computing can be sold the same way.
The claim is older than most people expect. Speaking at MIT's centennial in 1961, John McCarthy predicted that computing may someday be organized as a public utility, just as the telephone system is. He imagined computer service companies with subscribers connected to them, where each subscriber pays only for the capacity actually used, yet has access to everything a very large system offers. That is Utility computing, and the slide's verdict is right: it is remarkably close to what we now call cloud computing.
Three parts of one sentence
Read McCarthy's sentence slowly and it splits into three ideas. First, there is shared central capacity: one very large system that many subscribers draw on. Second, usage is metered and billed, so you pay for what you consume rather than for what you own. Third, subscribers reach the system over a network, from wherever they are. Fifty years later, NIST SP 800-145 wrote the formal definition of cloud computing, and three of its five essential characteristics line up with these ideas almost word for word.
| McCarthy's words | What it means | NIST characteristic |
|---|---|---|
| Access to all programming languages characteristic of a very large system | Many subscribers share one big pool | Resource pooling |
| Pay only for the capacity actually used | Usage is metered and billed | Measured service |
| Subscribers connected to a computer service company | Reached over a network from anywhere | Broad network access |
Part 06 teaches all five NIST characteristics and the deployment models. For now, keep the pattern: cloud computing is first an economic arrangement, and the technology exists to make that arrangement safe and cheap.
Recall
Which three parts of McCarthy's 1961 vision line up with modern cloud characteristics?
Quick check
McCarthy's "pay only for the capacity that he actually uses" anticipates which NIST characteristic most directly?
Picture a company that runs its ERP system on premise. It leases a building and fits it with power and cooling, buys racks, switches, a storage array and servers, installs an operating system, runs a database, patches security holes and finally deploys the application its staff actually use. Every one of those steps needs its own people and its own budget line, and nobody else is responsible if any of them fails.
That is Legacy IT. The defining rule is simple: the customer manages every layer. The slide draws nine of them, and they are worth learning bottom to top, because every cloud service model in part 07 is defined by how many of these same layers move to a provider. Infrastructure as a service takes the bottom ones, platform as a service takes more, and software as a service takes them all.
Legacy IT layers, bottom to top, and what the customer must do
- Data center
- Provide the building, power, cooling and physical security that host applications and data.
- Networking
- Buy and run the switches, routers and cabling that physically connect the servers.
- Storage
- Provide disks and arrays that hold large amounts of data, plus backup.
- Server
- Buy, rack, repair and refresh the machines that run the business applications.
- Virtualization
- Optional: run a hypervisor to share servers. Classic legacy IT had none (see the note below).
- Operating system
- Install, license, patch and upgrade the OS on every machine.
- Databases
- Install, tune, back up and replicate the database engines.
- Security
- Run firewalls, access control, monitoring and patching across the stack.
- Applications
- Deploy, configure and maintain the software the users actually need.
The slide's four bullets name the expensive physical base of this stack: data center technology to operate applications and store data, servers to run the business applications, networking to connect the servers, and storage for large amounts of data. Those are exactly the layers that part 05 will price as capital expenditure, and they are the first layers a cloud provider takes over. NIST SP 800-145 phrases each service model the same way, by listing which layers the consumer still controls.
Recall
In legacy IT, which layers does the customer manage?
Quick check
In the legacy IT model, who manages the operating system layer?
Take a server with 64 cores and 256 GB of RAM that runs one web application at about 10% load. Most of that machine is idle most of the time, yet it still uses power, rack space and an administrator's attention. Now carve it into six virtual machines: a Linux web server, a Linux database, a Windows directory server and three more. Each one gets its own virtual CPUs, its own slice of RAM, its own virtual disk and its own network card, and each boots its own operating system.
That carving is Virtualization: the resources of one computer are split so that several independent operating system instances can run on it at the same time. Each piece is a Virtual machine (VM), and it has two defining traits. It behaves like any other computer, with its own components, and it runs inside an isolated environment on the physical machine. NIST SP 800-125 puts it formally: virtualization is the simulation of the software and hardware on which other software runs, and that simulated environment is called a virtual machine. The guest appears to have its own processor, memory, storage controllers, Ethernet controllers, display, keyboard and mouse.
Two payoffs: consolidation and isolation
The first payoff is consolidation. NIST calls it operational efficiency: you place more load on each physical computer, so you buy, power and cool fewer of them. The hardware stops being a set of fixed boxes and becomes a pool, which is exactly the Resource pooling that McCarthy's utility needed.
The second payoff is isolation. Smith and Nair state it plainly: if security on one guest system is compromised, or one guest operating system fails, the software on the other guests is not affected. NIST adds that the hypervisor partitions the system's resources so that each guest can reach only its own. Isolation is what lets a provider rent neighbouring VMs to strangers, and it is the "safely" in the slide's summary that virtualization turns hardware into a pool that can safely run many independent computers.
Worked example
Consolidating five lightly loaded servers (illustrative numbers)
State the starting point
Five physical servers of identical size run at 10%, 12%, 15%, 11% and 14% average CPU utilization. These figures are made up for teaching, but single-digit to low-teen utilization is the typical shape of one-application-per-server estates.Add the loads as if they shared one host
10 + 12 + 15 + 11 + 14 = 62% of one server of the same size.Leave headroom
The combined load is well under 100%. The remaining 38% absorbs peaks that do not coincide and the small overhead of the hypervisor itself.Result
One host replaces five, each workload keeps its own OS and isolation, and four machines' worth of power, space and maintenance disappears. If peaks did line up, you would size for the peak instead of the average.
Recall
What are the two payoffs of virtualization, and what does each one buy a cloud provider?
Install Ubuntu in VirtualBox on an x86 laptop. Ubuntu's own code runs directly on the real CPU at full speed. Only when Ubuntu tries something sensitive, such as changing page tables or touching a device, does the virtualization software step in, perform the operation on Ubuntu's behalf against the real hardware, and hand back the result. Ubuntu never notices. Now run an old console game in a Nintendo emulator: every guest instruction has to be translated, because the guest speaks a different instruction set. (BlueStacks, the slide's other example, does the same for the ARM code inside Android apps.) Finally, sit in a flight simulator. It does not run the aircraft's software at all; it models how the aircraft behaves.
These three scenes are virtualization, emulation and Simulation. The first property to name is transparency. Inside a VM you install an operating system and applications exactly as on a physical computer, and the applications do not notice they are in a VM. Requests from the guest are transparently intercepted by the Virtual machine monitor (VMM) and converted for the real hardware, and the VM never becomes aware of the layer between it and the machine. Smith and Nair describe the mechanism: when a guest performs a privileged instruction, the VMM intercepts it, checks it and performs it on behalf of the guest, and guest software is unaware of this behind-the-scenes work.
Popek and Goldberg: what counts as a VM
In 1974, Popek and Goldberg gave the precise version. A virtual machine is "an efficient, isolated duplicate of the real machine", and a VMM must have three properties.
- Equivalence. Programs see an environment essentially identical to the original machine. Apart from timing and resource availability, they behave as they would on bare hardware.
- Efficiency. A statistically dominant subset of the guest's instructions runs directly on the real processor, with no software intervention.
- Resource control. The VMM stays in complete control of the real resources. A guest cannot reach resources it was not allocated.
The efficiency property is the one that draws the line. Popek and Goldberg say explicitly that it rules out traditional emulators and complete software interpreters (simulators) from the virtual machine umbrella. Their Theorem 1 gives a sufficient condition for building a VMM this way: every sensitive instruction is also privileged, so it traps to the VMM instead of running silently.
| Aspect | Virtualization | Emulation | Simulation |
|---|---|---|---|
| Core idea | Creates logical copies of physical resources | Mimics a hardware/software interface | Models a system's behaviour |
| What runs natively | Most guest instructions | Nothing: every instruction is translated | The model, not the real software |
| Same ISA required | Yes | No, that is the point | Not applicable |
| Guest aware? | No, interception is transparent | No, the translation is hidden | There is no guest |
| Speed | Close to native | Much slower than native | Whatever the model needs |
| Use case | Running many OSs or VMs on one host | Running software built for other hardware | Training or testing |
| Example | VMware, VirtualBox | BlueStacks, Nintendo emulator | Flight simulator, VR training |
Emulation still matters. Smith and Nair note that a whole-system VM whose ISA differs from the host, such as Virtual PC running Windows on a PowerPC Mac, must emulate both application and OS code. It works, and it is how you run old consoles or foreign architectures, but you pay for every instruction.
Recall
State Popek and Goldberg's three VMM properties, and the one that excludes emulators.
Quick check
Software interprets almost every guest instruction because the guest uses a different ISA. By Popek and Goldberg's definition, what is it?
In 1972 an IBM mainframe could be shared by hundreds of users at once. The trick was the Control Program (CP), which gave each user a complete virtual machine, a virtual System/370 duplicating the physical hardware. Inside that VM each user ran CMS, a simple single-user operating system. Many single-user machines, each in its own VM, added up to a multi-user time-sharing system. That product was VM/370, and Creasy's 1981 paper cited on the slide tells its story.
The lineage starts earlier. CP/CMS was conceived in 1964, CP-40 and CMS were running in 1966, and CP-67 on the System/360 Model 67 was in production use from 1967. IBM announced VM/370 on August 2, 1972. Its descendant, z/VM, still runs on IBM mainframes, so virtualization is one of the longest-lived ideas in computing.
A family, classified by what is virtualized
Today "virtualization" names a family of techniques. They differ in which layer is virtualized and how much of the guest is changed.
- Partitioning splits one machine's hardware into fixed, separate partitions.
- Hardware emulation presents hardware different from the real platform, at a translation cost.
- Application virtualization wraps one application so it runs isolated from the host OS.
- Full virtualization presents virtual hardware close enough to the real thing that unmodified guest OSs run.
- Paravirtualization modifies the guest kernel to call the hypervisor directly instead of touching emulated hardware.
- Hardware virtualization gives each complete guest OS its own virtual hardware, which the VMM maps onto shared physical CPU, memory, storage and network (slide 21). CPU extensions such as Intel VT-x andAMD-V (hardware-assisted virtualization) make this efficient.
- OS-level virtualization shares one kernel and isolates user spaces, which gives containers.
- Storage virtualization pools disks behind a SAN or software-defined storage.
- Network virtualization carves logical networks out of shared links with VLAN or VXLAN.
Hardware assistance exists because x86 does not meet Theorem 1's condition. Barham and colleagues note that some x86 supervisor instructions fail silently instead of trapping, so a VMM cannot catch them. VMware ESX Server solved this for full virtualization by dynamically rewriting guest code, Xen sidestepped it with paravirtualization, and later CPUs added extensions that make the instructions trap.
Two independent axes
The most important correction in this part is that two different questions get mixed up. The first axis is guest modification: does the guest OS run unmodified (full virtualization), or is its kernel changed to cooperate with the hypervisor (paravirtualization)? Barham and colleagues define paravirtualization exactly this way: a virtual machine abstraction similar but not identical to the hardware, which requires modifications to the guest OS but none to guest applications.
The second axis is hypervisor placement: does the VMM run directly on the hardware (Type 1 hypervisor) or on top of a host operating system (Type 2 hypervisor)? Goldberg's 1973 thesis defines it so: a Type I VMM runs on a hardware host, a Type II VMM runs on an extended host, meaning an operating system. Part 04 teaches the two types in depth. NIST SP 800-125 confirms the axes are independent by naming bare metal and hosted as the two forms of full virtualization.
| Unmodified guest (full) | Modified guest (para) | |
|---|---|---|
| Type 1, bare metal | VMware ESX/ESXi, Xen HVM | Xen PV |
| Type 2, hosted | VirtualBox, VMware Workstation | Paravirtual storage and network drivers in hosted products |
Recall
Why is "paravirtualization = Type 1" wrong?
Quick check
Which statement correctly fixes slide 20's mapping of virtualization techniques to hypervisor types?
Three VMs share one host. Follow a single disk write from one of them. The application calls a library function to save a file. The library asks the guest OS. The guest OS drives what it believes is its disk controller, but that controller is virtual hardware. When the guest touches it, the access reaches the hardware/software interface, where the Virtual machine monitor (VMM) catches it and maps it onto the real SSD that all three VMs share.
Saves a file.
Calls the OS.
Drives its disk driver.
A virtual disk controller.
Maps virtual to physical.
CPU, memory, storage, network.
The rule behind the picture is a single horizontal cut. Everything above the hardware/software interface, the application, library, operating system and virtual hardware, is replicated for each VM and isolated from the others. Everything below it, the CPU, memory, storage and network, is physical, shared and multiplexed by the VMM. Smith and Nair say the VMM has access to, and manages, all the hardware resources, and NIST says the hypervisor controls the flow of instructions between the guests and the physical CPU, disk, memory and network cards. So the OS instances are isolated from each other, and each runs on virtualized resources that are mapped to physical ones.
Worked example
Tracing one disk write through the layers
Application and library
The program calls a file write. The C library turns it into a system call to the guest kernel. Nothing special has happened yet: this all runs natively on the real CPU.Guest OS
The guest kernel's file system and disk driver build a request and write it to the registers of its disk controller, exactly as they would on a physical machine.Interception at the interface
Those controller registers are virtual. The access traps to the VMM, which checks that the request touches only this VM's virtual disk.Shared hardware
The VMM translates the virtual block address to a location in the VM's disk image on the shared SSD, issues the real I/O, and later signals completion back into the guest as a virtual interrupt.Result
The guest saw an ordinary disk write. Isolation held because the VMM only ever mapped it onto this VM's share of the device.
Where you cut decides the kind of virtualization
The interface on the slide is essentially the instruction set architecture. Smith and Nair write that the ISA marks the division between hardware and software, and they identify three key interfaces: the ISA, the ABI (system calls plus user instructions) and the API (library calls). Cut the stack at the ISA and you get a system VM with a full guest OS, as on this slide. Cut higher, at the ABI or API, and you get process VMs or OS-level containers, where the kernel is shared and only the user space is isolated. That is the bridge to part 04.
Recall
In the hardware virtualization stack, what is isolated and what is shared?
Recap
If you remember nothing else
- McCarthy (1961) described computing sold like a utility: shared capacity, pay per use, network access. That is the economic idea behind cloud computing.
- In legacy IT the customer runs every layer, from the data center to applications. Cloud models move layers to a provider.
- A VM is an efficient, isolated duplicate of a real machine. The guest OS does not notice the VMM intercepting sensitive operations.
- Virtualization runs most instructions natively on the same ISA. Emulation translates between ISAs. Simulation models behaviour without running the real software.
- Virtualization began with IBM CP-40 and CP-67, and VM/370 (announced 1972) turned time-sharing into one VM per user.
- Full versus para (is the guest modified?) and Type 1 versus Type 2 (where does the hypervisor run?) are independent axes.
- Above the hardware/software interface everything is isolated per VM. Below it, the physical hardware is shared.
Sources
- Formal Requirements for Virtualizable Third Generation ArchitecturesPaperPopek and Goldberg, Communications of the ACM 17(7), 1974Equivalence, efficiency, resource control, and Theorem 1.(opens in a new tab)
- The Origin of the VM/370 Time-sharing SystemPaperCreasy, IBM Journal of Research and Development 25(5), 1981CP-40, CP-67 and VM/370 history.(opens in a new tab)
- Xen and the Art of VirtualizationPaperBarham et al., SOSP 2003Definition of paravirtualization and the x86 trap problem.(opens in a new tab)
- The Architecture of Virtual MachinesPaperSmith and Nair, IEEE Computer 38(5), 2005ISA, ABI and API interfaces; isolation; process versus system VMs.(opens in a new tab)
- Architectural Principles for Virtual Computer SystemsPaperGoldberg, Harvard PhD thesis, 1973Origin of the Type I and Type II VMM distinction.(opens in a new tab)
- SP 800-125: Guide to Security for Full Virtualization TechnologiesDocsNIST, 2011Definition of a VM, bare metal and hosted full virtualization, hardware emulation.(opens in a new tab)
- SP 800-145: The NIST Definition of Cloud ComputingDocsNIST, 2011(opens in a new tab)
- z/VM History: TimelineDocsIBMVM/370 announced August 2, 1972.(opens in a new tab)
- The Cloud ImperativeArticleGarfinkel, MIT Technology Review, 2011McCarthy's 1961 utility computing quote.(opens in a new tab)
- Architects of the Information SocietyBookGarfinkel, MIT Press, 1999(opens in a new tab)
- Understanding Full Virtualization, Paravirtualization, and Hardware AssistArticleVMware, 2007Further reading.(opens in a new tab)