Majid Al-RaimiFrom legacy IT to virtualization

COE 558Lecture 02Part 03

From legacy IT to virtualization

Starts the cloud computing section with the utility computing vision and the legacy IT stack, then defines virtualization, its history and its main forms.

Concepts
6
Slides
15-21
Reading
36 min
Understood
0/6 concepts

Why this part matters

This part opens the cloud half of the lecture. Every cloud product you will price in parts 05 and 06, classify in part 07 or deploy in parts 08 and 09 is, at heart, a way of handing some layers of the classic IT stack to someone else. Before you can say which layers move, you need to know what the stack is and which technology lets one provider sell slices of a machine to many customers at once.

That technology is virtualization, and it is what turned a 1961 economic idea into a business. For your research and for the edge systems you will design, the trade-offs between a virtual machine, an emulator and a container decide how much isolation you get, how much overhead you pay and how portable your workload is. This part builds those distinctions carefully, and it flags two places where the slides blur them.

By the end you can

  1. Explain McCarthy's utility computing vision and map it to modern cloud characteristics.
  2. List the layers of legacy IT and say who manages each one.
  3. Define a virtual machine and explain transparency using Popek and Goldberg's three properties.
  4. Distinguish virtualization, emulation and simulation with an example of each.
  5. Classify virtualization techniques on two independent axes: guest modification and hypervisor placement.
  6. Explain which resources are isolated and which are shared in the hardware virtualization stack.

Look at your electricity bill. You never bought a power station, you never hired the engineers who run it, and you can plug in a kettle or a server without asking anyone. You pay for the kilowatt-hours you actually used, and the grid is reached through a socket in the wall. Computing can be sold the same way.

The claim is older than most people expect. Speaking at MIT's centennial in 1961, John McCarthy predicted that computing may someday be organized as a public utility, just as the telephone system is. He imagined computer service companies with subscribers connected to them, where each subscriber pays only for the capacity actually used, yet has access to everything a very large system offers. That is Utility computing, and the slide's verdict is right: it is remarkably close to what we now call cloud computing.

A socket, a cable and a meter: the dial and the counter move only while compute flows, which is pay per use.

Three parts of one sentence

Read McCarthy's sentence slowly and it splits into three ideas. First, there is shared central capacity: one very large system that many subscribers draw on. Second, usage is metered and billed, so you pay for what you consume rather than for what you own. Third, subscribers reach the system over a network, from wherever they are. Fifty years later, NIST SP 800-145 wrote the formal definition of cloud computing, and three of its five essential characteristics line up with these ideas almost word for word.

McCarthy's wordsWhat it meansNIST characteristic
Access to all programming languages characteristic of a very large systemMany subscribers share one big poolResource pooling
Pay only for the capacity actually usedUsage is metered and billedMeasured service
Subscribers connected to a computer service companyReached over a network from anywhereBroad network access
McCarthy (1961) against NIST SP 800-145 (2011)

Part 06 teaches all five NIST characteristics and the deployment models. For now, keep the pattern: cloud computing is first an economic arrangement, and the technology exists to make that arrangement safe and cheap.

Recall

Which three parts of McCarthy's 1961 vision line up with modern cloud characteristics?

Shared central capacity (resource pooling), paying only for the capacity used (measured service), and subscribers connected to a service company (broad network access).

Quick check

McCarthy's "pay only for the capacity that he actually uses" anticipates which NIST characteristic most directly?

Picture a company that runs its ERP system on premise. It leases a building and fits it with power and cooling, buys racks, switches, a storage array and servers, installs an operating system, runs a database, patches security holes and finally deploys the application its staff actually use. Every one of those steps needs its own people and its own budget line, and nobody else is responsible if any of them fails.

That is Legacy IT. The defining rule is simple: the customer manages every layer. The slide draws nine of them, and they are worth learning bottom to top, because every cloud service model in part 07 is defined by how many of these same layers move to a provider. Infrastructure as a service takes the bottom ones, platform as a service takes more, and software as a service takes them all.

Nine layers, one owner: a bracket runs the full height and every bar lights for the customer, while the virtualization layer stays in question.

Legacy IT layers, bottom to top, and what the customer must do

Data center
Provide the building, power, cooling and physical security that host applications and data.
Networking
Buy and run the switches, routers and cabling that physically connect the servers.
Storage
Provide disks and arrays that hold large amounts of data, plus backup.
Server
Buy, rack, repair and refresh the machines that run the business applications.
Virtualization
Optional: run a hypervisor to share servers. Classic legacy IT had none (see the note below).
Operating system
Install, license, patch and upgrade the OS on every machine.
Databases
Install, tune, back up and replicate the database engines.
Security
Run firewalls, access control, monitoring and patching across the stack.
Applications
Deploy, configure and maintain the software the users actually need.

The slide's four bullets name the expensive physical base of this stack: data center technology to operate applications and store data, servers to run the business applications, networking to connect the servers, and storage for large amounts of data. Those are exactly the layers that part 05 will price as capital expenditure, and they are the first layers a cloud provider takes over. NIST SP 800-145 phrases each service model the same way, by listing which layers the consumer still controls.

Recall

In legacy IT, which layers does the customer manage?

All of them: data center, networking, storage, server, virtualization (if any), operating system, databases, security and applications.

Quick check

In the legacy IT model, who manages the operating system layer?

Take a server with 64 cores and 256 GB of RAM that runs one web application at about 10% load. Most of that machine is idle most of the time, yet it still uses power, rack space and an administrator's attention. Now carve it into six virtual machines: a Linux web server, a Linux database, a Windows directory server and three more. Each one gets its own virtual CPUs, its own slice of RAM, its own virtual disk and its own network card, and each boots its own operating system.

That carving is Virtualization: the resources of one computer are split so that several independent operating system instances can run on it at the same time. Each piece is a Virtual machine (VM), and it has two defining traits. It behaves like any other computer, with its own components, and it runs inside an isolated environment on the physical machine. NIST SP 800-125 puts it formally: virtualization is the simulation of the software and hardware on which other software runs, and that simulated environment is called a virtual machine. The guest appears to have its own processor, memory, storage controllers, Ethernet controllers, display, keyboard and mouse.

Three VMs rise from one hardware slab, each with its own guest OS. A request from one VM is caught at the VMM line and routed into the shared hardware.

Two payoffs: consolidation and isolation

The first payoff is consolidation. NIST calls it operational efficiency: you place more load on each physical computer, so you buy, power and cool fewer of them. The hardware stops being a set of fixed boxes and becomes a pool, which is exactly the Resource pooling that McCarthy's utility needed.

The second payoff is isolation. Smith and Nair state it plainly: if security on one guest system is compromised, or one guest operating system fails, the software on the other guests is not affected. NIST adds that the hypervisor partitions the system's resources so that each guest can reach only its own. Isolation is what lets a provider rent neighbouring VMs to strangers, and it is the "safely" in the slide's summary that virtualization turns hardware into a pool that can safely run many independent computers.

Worked example

Consolidating five lightly loaded servers (illustrative numbers)

  1. State the starting point

    Five physical servers of identical size run at 10%, 12%, 15%, 11% and 14% average CPU utilization. These figures are made up for teaching, but single-digit to low-teen utilization is the typical shape of one-application-per-server estates.
  2. Add the loads as if they shared one host

    10 + 12 + 15 + 11 + 14 = 62% of one server of the same size.
  3. Leave headroom

    The combined load is well under 100%. The remaining 38% absorbs peaks that do not coincide and the small overhead of the hypervisor itself.
  4. Result

    One host replaces five, each workload keeps its own OS and isolation, and four machines' worth of power, space and maintenance disappears. If peaks did line up, you would size for the peak instead of the average.

Recall

What are the two payoffs of virtualization, and what does each one buy a cloud provider?

Consolidation (more load per physical host, so fewer machines to buy, power and cool, and hardware becomes a pool) and isolation (a compromised or crashed guest does not affect the others, so strangers can rent neighbouring VMs).

Install Ubuntu in VirtualBox on an x86 laptop. Ubuntu's own code runs directly on the real CPU at full speed. Only when Ubuntu tries something sensitive, such as changing page tables or touching a device, does the virtualization software step in, perform the operation on Ubuntu's behalf against the real hardware, and hand back the result. Ubuntu never notices. Now run an old console game in a Nintendo emulator: every guest instruction has to be translated, because the guest speaks a different instruction set. (BlueStacks, the slide's other example, does the same for the ARM code inside Android apps.) Finally, sit in a flight simulator. It does not run the aircraft's software at all; it models how the aircraft behaves.

These three scenes are virtualization, emulation and Simulation. The first property to name is transparency. Inside a VM you install an operating system and applications exactly as on a physical computer, and the applications do not notice they are in a VM. Requests from the guest are transparently intercepted by the Virtual machine monitor (VMM) and converted for the real hardware, and the VM never becomes aware of the layer between it and the machine. Smith and Nair describe the mechanism: when a guest performs a privileged instruction, the VMM intercepts it, checks it and performs it on behalf of the guest, and guest software is unaware of this behind-the-scenes work.

Popek and Goldberg: what counts as a VM

In 1974, Popek and Goldberg gave the precise version. A virtual machine is "an efficient, isolated duplicate of the real machine", and a VMM must have three properties.

  1. Equivalence. Programs see an environment essentially identical to the original machine. Apart from timing and resource availability, they behave as they would on bare hardware.
  2. Efficiency. A statistically dominant subset of the guest's instructions runs directly on the real processor, with no software intervention.
  3. Resource control. The VMM stays in complete control of the real resources. A guest cannot reach resources it was not allocated.

The efficiency property is the one that draws the line. Popek and Goldberg say explicitly that it rules out traditional emulators and complete software interpreters (simulators) from the virtual machine umbrella. Their Theorem 1 gives a sufficient condition for building a VMM this way: every sensitive instruction is also privileged, so it traps to the VMM instead of running silently.

Sensitive⊆Privileged\text{Sensitive} \subseteq \text{Privileged}
Popek and Goldberg Theorem 1
AspectVirtualizationEmulationSimulation
Core ideaCreates logical copies of physical resourcesMimics a hardware/software interfaceModels a system's behaviour
What runs nativelyMost guest instructionsNothing: every instruction is translatedThe model, not the real software
Same ISA requiredYesNo, that is the pointNot applicable
Guest aware?No, interception is transparentNo, the translation is hiddenThere is no guest
SpeedClose to nativeMuch slower than nativeWhatever the model needs
Use caseRunning many OSs or VMs on one hostRunning software built for other hardwareTraining or testing
ExampleVMware, VirtualBoxBlueStacks, Nintendo emulatorFlight simulator, VR training
Virtualization, emulation and simulation

Emulation still matters. Smith and Nair note that a whole-system VM whose ISA differs from the host, such as Virtual PC running Windows on a PowerPC Mac, must emulate both application and OS code. It works, and it is how you run old consoles or foreign architectures, but you pay for every instruction.

Recall

State Popek and Goldberg's three VMM properties, and the one that excludes emulators.

Equivalence, efficiency and resource control. Efficiency excludes emulators and simulators, because a statistically dominant subset of instructions must run directly on the real processor.

Quick check

Software interprets almost every guest instruction because the guest uses a different ISA. By Popek and Goldberg's definition, what is it?

In 1972 an IBM mainframe could be shared by hundreds of users at once. The trick was the Control Program (CP), which gave each user a complete virtual machine, a virtual System/370 duplicating the physical hardware. Inside that VM each user ran CMS, a simple single-user operating system. Many single-user machines, each in its own VM, added up to a multi-user time-sharing system. That product was VM/370, and Creasy's 1981 paper cited on the slide tells its story.

The lineage starts earlier. CP/CMS was conceived in 1964, CP-40 and CMS were running in 1966, and CP-67 on the System/360 Model 67 was in production use from 1967. IBM announced VM/370 on August 2, 1972. Its descendant, z/VM, still runs on IBM mainframes, so virtualization is one of the longest-lived ideas in computing.

A family, classified by what is virtualized

Today "virtualization" names a family of techniques. They differ in which layer is virtualized and how much of the guest is changed.

  • Partitioning splits one machine's hardware into fixed, separate partitions.
  • Hardware emulation presents hardware different from the real platform, at a translation cost.
  • Application virtualization wraps one application so it runs isolated from the host OS.
  • Full virtualization presents virtual hardware close enough to the real thing that unmodified guest OSs run.
  • Paravirtualization modifies the guest kernel to call the hypervisor directly instead of touching emulated hardware.
  • Hardware virtualization gives each complete guest OS its own virtual hardware, which the VMM maps onto shared physical CPU, memory, storage and network (slide 21). CPU extensions such as Intel VT-x andAMD-V (hardware-assisted virtualization) make this efficient.
  • OS-level virtualization shares one kernel and isolates user spaces, which gives containers.
  • Storage virtualization pools disks behind a SAN or software-defined storage.
  • Network virtualization carves logical networks out of shared links with VLAN or VXLAN.

Hardware assistance exists because x86 does not meet Theorem 1's condition. Barham and colleagues note that some x86 supervisor instructions fail silently instead of trapping, so a VMM cannot catch them. VMware ESX Server solved this for full virtualization by dynamically rewriting guest code, Xen sidestepped it with paravirtualization, and later CPUs added extensions that make the instructions trap.

Two independent axes

The most important correction in this part is that two different questions get mixed up. The first axis is guest modification: does the guest OS run unmodified (full virtualization), or is its kernel changed to cooperate with the hypervisor (paravirtualization)? Barham and colleagues define paravirtualization exactly this way: a virtual machine abstraction similar but not identical to the hardware, which requires modifications to the guest OS but none to guest applications.

The second axis is hypervisor placement: does the VMM run directly on the hardware (Type 1 hypervisor) or on top of a host operating system (Type 2 hypervisor)? Goldberg's 1973 thesis defines it so: a Type I VMM runs on a hardware host, a Type II VMM runs on an extended host, meaning an operating system. Part 04 teaches the two types in depth. NIST SP 800-125 confirms the axes are independent by naming bare metal and hosted as the two forms of full virtualization.

Unmodified guest (full)Modified guest (para)
Type 1, bare metalVMware ESX/ESXi, Xen HVMXen PV
Type 2, hostedVirtualBox, VMware WorkstationParavirtual storage and network drivers in hosted products
Placement (rows) by guest modification (columns)
At rest, the slide's diagonal pairing (full with Type 2, para with Type 1). When active, all four quadrants fill: the axes are independent.

Recall

Why is "paravirtualization = Type 1" wrong?

Paravirtualization means the guest OS is modified to cooperate with the hypervisor. Type 1 versus Type 2 is about whether the hypervisor runs on bare metal or on a host OS. These are independent axes: ESXi does full virtualization on bare metal, and hosted products such as VirtualBox and VMware Workstation use paravirtual storage and network drivers (NIST SP 800-125 notes this paravirtualization of storage and Ethernet controllers).

Quick check

Which statement correctly fixes slide 20's mapping of virtualization techniques to hypervisor types?

Three VMs share one host. Follow a single disk write from one of them. The application calls a library function to save a file. The library asks the guest OS. The guest OS drives what it believes is its disk controller, but that controller is virtual hardware. When the guest touches it, the access reaches the hardware/software interface, where the Virtual machine monitor (VMM) catches it and maps it onto the real SSD that all three VMs share.

Application
isolated

Saves a file.

Library
isolated

Calls the OS.

Guest OS
isolated

Drives its disk driver.

Virtual hardware
isolated

A virtual disk controller.

trap
HSI
VMM intercepts

Maps virtual to physical.

Hardware
shared

CPU, memory, storage, network.

The hardware virtualization stack, from the application down to the shared hardware

The rule behind the picture is a single horizontal cut. Everything above the hardware/software interface, the application, library, operating system and virtual hardware, is replicated for each VM and isolated from the others. Everything below it, the CPU, memory, storage and network, is physical, shared and multiplexed by the VMM. Smith and Nair say the VMM has access to, and manages, all the hardware resources, and NIST says the hypervisor controls the flow of instructions between the guests and the physical CPU, disk, memory and network cards. So the OS instances are isolated from each other, and each runs on virtualized resources that are mapped to physical ones.

Worked example

Tracing one disk write through the layers

  1. Application and library

    The program calls a file write. The C library turns it into a system call to the guest kernel. Nothing special has happened yet: this all runs natively on the real CPU.
  2. Guest OS

    The guest kernel's file system and disk driver build a request and write it to the registers of its disk controller, exactly as they would on a physical machine.
  3. Interception at the interface

    Those controller registers are virtual. The access traps to the VMM, which checks that the request touches only this VM's virtual disk.
  4. Shared hardware

    The VMM translates the virtual block address to a location in the VM's disk image on the shared SSD, issues the real I/O, and later signals completion back into the guest as a virtual interrupt.
  5. Result

    The guest saw an ordinary disk write. Isolation held because the VMM only ever mapped it onto this VM's share of the device.

Where you cut decides the kind of virtualization

The interface on the slide is essentially the instruction set architecture. Smith and Nair write that the ISA marks the division between hardware and software, and they identify three key interfaces: the ISA, the ABI (system calls plus user instructions) and the API (library calls). Cut the stack at the ISA and you get a system VM with a full guest OS, as on this slide. Cut higher, at the ABI or API, and you get process VMs or OS-level containers, where the kernel is shared and only the user space is isolated. That is the bridge to part 04.

Recall

In the hardware virtualization stack, what is isolated and what is shared?

Application, library, OS and virtual hardware are isolated per VM, above the hardware/software interface. CPU, memory, storage and networking are shared physical resources below it, multiplexed by the VMM.

Recap

If you remember nothing else

  • McCarthy (1961) described computing sold like a utility: shared capacity, pay per use, network access. That is the economic idea behind cloud computing.
  • In legacy IT the customer runs every layer, from the data center to applications. Cloud models move layers to a provider.
  • A VM is an efficient, isolated duplicate of a real machine. The guest OS does not notice the VMM intercepting sensitive operations.
  • Virtualization runs most instructions natively on the same ISA. Emulation translates between ISAs. Simulation models behaviour without running the real software.
  • Virtualization began with IBM CP-40 and CP-67, and VM/370 (announced 1972) turned time-sharing into one VM per user.
  • Full versus para (is the guest modified?) and Type 1 versus Type 2 (where does the hypervisor run?) are independent axes.
  • Above the hardware/software interface everything is isolated per VM. Below it, the physical hardware is shared.

Sources