COE 558Lecture 02Part 04
Hypervisors, containers, namespaces and cgroups
Compares Type 1 and Type 2 hypervisors, then explains OS-level virtualization (containers) built from Linux namespaces and control groups.
- Concepts
- 6
- Slides
- 22-28
- Reading
- 36 min
Why this part matters
Every cloud service in the rest of this course, from EC2 and Lambda to Kubernetes on EKS or GKE and edge nodes running k3s, is built from one of two isolation designs. Either a hypervisor splits the hardware into virtual machines, or one shared kernel is split into containers by namespaces and cgroups.
The choice between them decides overhead, density, start-up time and how strong the security boundary is. Comparing Type 1, Type 2 and containers, and telling namespaces from cgroups, are the likely exam questions. The same choice comes back when you size edge devices for the research project, where a small gateway may not have the memory to boot a full guest OS for every service.
By the end you can
- Place Type 1 and Type 2 hypervisors in the software stack and name real examples of each, including where KVM and Hyper-V fit.
- Compare Type 1 and Type 2 on overhead, resource management, administration and isolation, and explain why cloud providers choose Type 1.
- Explain OS-level virtualization: what containers share, what they isolate, and why their security boundary is weaker than a VM's.
- Distinguish namespaces (what a process sees) from cgroups (what a process uses), and name the main namespace types and cgroup functions.
- Read a cgroup CPU quota and weight and predict how a container behaves under contention and at its memory limit.
Start with a laptop running VirtualBox with Ubuntu inside it. When you press the power button, macOS boots first. VirtualBox is then just another application, next to Safari, and the whole Ubuntu guest lives inside that application's process. Now compare a server in an AWS data center that hosts EC2 instances. Nothing general purpose boots before the hypervisor. It is the first thing on the machine, and it hands out CPU and memory to the guests directly.
Both machines do Virtualization, and both run a Virtual machine monitor (VMM) that makes each Virtual machine (VM) believe it has a real computer. The difference is where the VMM sits, and that comes down to one question: who owns the hardware?
Two placements
- A Type 1 hypervisor (bare metal, or native) is the lowest software layer on the machine. It runs directly on the hardware and acts as a specialized OS whose only job is to run guests. Red Hat describes it as a hypervisor that "runs directly on the host's hardware to manage guest operating systems".
- A Type 2 hypervisor (hosted) runs on a conventional OS "as a software layer or application". The host OS owns the hardware, and the VMM is one of its processes, with less privilege than the OS itself. VirtualBox's own manual calls it "a so-called hosted hypervisor, sometimes referred to as a type 2 hypervisor".
Guest OS 1 to N run directly on the hypervisor, which owns CPU, memory, storage and network.
The VMM is an app on the host OS. Each guest request crosses the VMM and the host kernel.
Real examples, and where the labels get blurry
VMware's Type 1 hypervisor is ESXi, described as "the hypervisor [that] runs virtual machines"; vCenter Server is the central management service for many ESXi hosts, and vSphere is the name of the whole suite. On the hosted side sit VirtualBox, VMware Workstation and Parallels Desktop.
Picture a lab PC running VirtualBox for a coursework VM, next to an EC2 c7i instance. When the coursework VM compiles code, its virtual CPU is just a thread that competes with Chrome and the antivirus for time on the host scheduler. The EC2 instance gets CPUs and memory partitioned for it by a minimal hypervisor, with no desktop OS in the way.
That picture gives the whole comparison. Because a Type 1 hypervisor owns the hardware, it adds less overhead (the slide says about 5%), manages resources directly and isolates guests more strongly, but it needs more skill to install and operate. A Type 2 hypervisor reaches the hardware only through the host OS, so it costs more (about 10%), controls resources indirectly and isolates less, but anyone can install it.
| Property | Type 1 (bare metal) | Type 2 (hosted) |
|---|---|---|
| Runs on | Directly on the hardware, as the lowest software layer | As a process of a general purpose host OS |
| Approx. overhead | ≈5% | ≈10% |
| Resource management | Direct: the hypervisor schedules CPUs and partitions memory itself | Indirect: every request goes through the host OS scheduler and drivers |
| Administration | Needs more skill: a dedicated server, remote management, clusters | Easy: install it like any desktop app |
| Isolation | Stronger: only a small hypervisor sits between tenants | Weaker: a compromise of the host OS exposes every guest |
| Typical users | Cloud providers and enterprise data centers | Developers, students, testers on laptops |
| Examples | VMware ESXi, Microsoft Hyper-V, KVM (debated), AWS Nitro | Oracle VirtualBox, VMware Workstation, Parallels Desktop |
Why cloud providers pick Type 1
A Cloud service provider (CSP) sells slices of the same physical server to strangers. Every percent of overhead is capacity it cannot sell, and every extra layer between tenants is attack surface. AWS shows where this leads. Its Nitro Hypervisor is "a custom-developed, minimized hypervisor based on KVM", cut down on purpose: it has "no networking stack, no general-purpose file system implementations, and no peripheral device driver support". Networking and storage are offloaded to dedicated Nitro Cards and exposed to guests through SR-IOV. The hypervisor does little more than partition CPU and memory, which is the Type 1 idea taken to its limit, and the design is lean enough that AWS can also sell bare metal instances.
Take the slide's own example: one server runs a database and a web server as plain processes. Two things can go wrong. First, the web server can list the database's processes, send them signals, read its files and bind its ports, so whoever compromises one has a view of the other. Second, a heavy query makes the database use every CPU core, and the website stalls.
These are different problems, and Linux fixes each with a different kernel feature. The first is about visibility. A Namespace gives a group of processes its own view of a kernel resource, so the web server cannot even see the database's processes, network stack or mount points. The second is about consumption. Control groups (cgroups) put quotas, weights and accounting on how much CPU, memory and I/O a group can use. Docker's security guide gives the same split: namespaces "provide the first and most straightforward form of isolation", while cgroups "implement resource accounting and limiting".
| Problem | Fix | Kernel feature | Question it answers |
|---|---|---|---|
| Processes see and touch each other (PIDs, ports, files) | Give each group its own view of kernel resources | Namespaces | What can I see? |
| Processes compete for the same CPU, memory and I/O | Put quotas, weights and accounting on resource use | Control groups (cgroups) | How much can I use? |
Neither feature is new. The first namespace (mount) arrived in Linux 2.4.19 in 2002, most of the rest landed between 2.6.15 and 2.6.26 (July 2008), and user namespaces were only completed in 3.8 (2013), which made unprivileged containers practical. Cgroups were started in 2006 and merged in 2.6.24. A Container is simply these two combined, plus a root filesystem image. NIST describes the second half with a concrete case: a host with 10 GB of memory can give 1 GB to each of nine containers, so no single one can starve the rest.
You can make a Namespace by hand on any Linux machine. The command below starts a shell in a new PID namespace and mounts a fresh /proc for it. Inside, ps a shows only two processes, and the shell believes it is PID 1.
sudo unshare --fork --pid --mount-proc bash
ps aPID TTY STAT TIME COMMAND
1 pts/0 S 0:00 bash
8 pts/0 R+ 0:00 ps aFrom a second terminal on the host, the same bash appears in ps -A under a large global PID such as 24517. Nothing was copied and no second kernel was started. The kernel simply shows that process a different numbering. NIST gives the container version: ps -A inside an Apache Container lists only httpd.
The rule
The Linux manual defines a namespace as something that "wraps a global system resource in an abstraction that makes it appear to the processes within the namespace that they have their own isolated instance". There is one kernel and many views. Every process belongs to exactly one namespace of each type, and a container runtime such as Docker creates a fresh set (all but user and time, by default) for every container. You can see a process's memberships as symbolic links in /proc/<pid>/ns/ (since Linux 3.8), and a tool can join an existing namespace with setns(2), which is how docker exec gets a shell into a running container.
| Namespace | Flag | What it isolates |
|---|---|---|
| Cgroup | CLONE_NEWCGROUP | The cgroup root directory a process sees |
| IPC | CLONE_NEWIPC | System V IPC objects and POSIX message queues |
| Network | CLONE_NEWNET | Network devices, IP stacks, routing tables, ports |
| Mount | CLONE_NEWNS | Mount points, so each container has its own filesystem tree |
| PID | CLONE_NEWPID | Process ID numbers, starting again at 1 |
| Time | CLONE_NEWTIME | Boot and monotonic clock offsets |
| User | CLONE_NEWUSER | User and group IDs, so root inside can be unprivileged outside |
| UTS | CLONE_NEWUTS | Hostname and NIS domain name |
The network namespace is what gives each container "its own networking stack, including unique interfaces and IP addresses", in NIST's words, so two containers can both listen on port 80. The user namespace, when enabled, lets a process be root inside its container while mapping to an unprivileged user on the host.
On a host with 2 CPUs, start MongoDB with docker run --cpus=1.5 --memory=512m mongo. Docker writes two numbers into the container's cgroup: a CPU quota of 150000 µs per 100000 µs period, and a memory ceiling of 512 MiB. If MongoDB grows past that ceiling and the kernel cannot reclaim memory, the OOM killer kills a process inside that cgroup only. The web server next to it never notices.
docker run -d --name db --cpus=1.5 --memory=512m mongo
cat /sys/fs/cgroup/system.slice/docker-$(docker inspect -f '{{.Id}}' db).scope/cpu.max
cat /sys/fs/cgroup/system.slice/docker-$(docker inspect -f '{{.Id}}' db).scope/memory.max150000 100000
536870912How cgroups work
The manual says cgroups "allow processes to be organized into hierarchical groups whose usage of various types of resources can then be limited and monitored". Each kind of resource is enforced by a controller, which Red Hat defines as "a kernel subsystem that represents a single resource, such as CPU time, memory, network bandwidth or disk I/O". The main controllers are cpu, memory, io, pids and cpuset; freezing, a separate freezer controller in cgroup v1, is the core file cgroup.freeze in v2. Cgroup v2, official since Linux 4.5, puts everything in one tree mounted at /sys/fs/cgroup, where every setting is a small file you can read and write.
The slide lists four functions, and each maps to concrete files:
The four cgroup functions and the cgroup v2 files behind them
- Limiting
- Hard caps: cpu.max, memory.max, io.max, pids.max
- Prioritization
- Shares under contention: cpu.weight, io.weight
- Accounting
- Measured use for monitoring and billing: cpu.stat, memory.current, io.stat
- Control
- Freezing and thawing a whole group: cgroup.freeze (the basis for checkpoint and restart with CRIU)
Worked example
Reading a CPU quota and a CPU weight
Turn cpu.max into CPUs
The file holds 150000 100000: a quota of CPU time, then the period it applies to. The format is $MAX $PERIOD, and the default is max 100000 (no limit).
Compare with the machine
On a 2-CPU host that is 1.5 / 2 = 75% of all CPU time. It is a hard ceiling: once the group has used 150 ms of CPU time in a 100 ms period, it waits for the next period, even if the machine is otherwise idle.
Add prioritization
Now two busy containers have --cpu-shares 1024 and 512 (cgroup v2 calls this cpu.weight, default 100, range 1 to 10000). When they compete, each gets its weight divided by the sum of weights:
This is a soft limit. If the second container goes idle, the first may use all the CPU it can get.
Result
Limits are ceilings that hold even on an idle machine; weights are shares that matter only under contention. That is how to read the slide's diagram where four groups each get 25%: with equal weights, each gets a quarter only while all four are busy.
Recall
In one sentence each, where does a Type 1 and a Type 2 hypervisor run, and who owns the hardware?
Recall
Why do cloud providers choose Type 1 hypervisors?
Recall
What is shared and what is isolated in OS-level virtualization?
Recall
Namespaces and cgroups: which question does each answer?
Recall
A container gets --cpus=1.5. What does cpu.max hold, and what does it mean?
Recall
Two busy containers have cpu.weight 200 and 100 and no cpu.max. What share does each get, and what happens when the second goes idle?
Quick check
A cloud provider wants minimal overhead and strong isolation between tenant VMs on each server. Which virtualization layer should it deploy?
Quick check
A database and a web server share one Linux host. The database starts using every CPU core and the website slows down. Which kernel feature fixes this?
Quick check
Inside a container, ps shows only two processes and your app has PID 1. What produces this view?
Quick check
Why can a Linux host run a Windows VM but not a native Windows container?
Recap
If you remember nothing else
- Type 1 hypervisors run on bare metal and act as the OS (ESXi, Hyper-V, KVM). Type 2 run as an app on a host OS (VirtualBox, VMware Workstation, Parallels).
- Type 1: about 5% overhead, direct resource control, stronger isolation, harder to run. Type 2: about 10% overhead, indirect control, easier, less isolation.
- Cloud providers use Type 1. AWS Nitro is a minimal KVM-based Type 1 hypervisor that offloads I/O to dedicated cards.
- Containers use no hypervisor. They share the host kernel below the system call interface and isolate only their user space.
- Namespaces control what a process can see: PID, network, mount, UTS, IPC, user, cgroup and time.
- Cgroups control what a process can use: limiting, prioritization, accounting and control (freezing a whole group, which CRIU builds on to checkpoint and restart).
- A container is several namespaces plus cgroups plus a root filesystem. In the cloud, containers usually run inside VMs.
Sources
- What is a hypervisor?DocsRed HatType 1 and Type 2 definitions and examples(opens in a new tab)
- What is KVM?DocsRed HatEach VM is a regular Linux process(opens in a new tab)
- Kernel Virtual MachineDocsKVM project(opens in a new tab)
- The Definitive KVM API DocumentationDocsLinux kernel(opens in a new tab)
- Hyper-V ArchitectureDocsMicrosoft LearnRoot and child partitions, VMBus, enlightened I/O(opens in a new tab)
- vSphere Software ComponentsDocsBroadcom TechDocsESXi is the hypervisor; vSphere is the suite(opens in a new tab)
- Oracle VirtualBox manual: IntroductionDocsOracle(opens in a new tab)
- Security Design of the AWS Nitro System: the componentsDocsAmazon Web Services(opens in a new tab)
- Security Design of the AWS Nitro System: the Nitro System journeyDocsAmazon Web Services(opens in a new tab)
- SP 800-190: Application Container Security GuidePaperNIST (Souppaya, Morello, Scarfone, 2017)Sections 2.2 and 3.5.2(opens in a new tab)
- namespaces(7)DocsLinux manual pages(opens in a new tab)
- unshare(1)DocsLinux manual pages(opens in a new tab)
- cgroups(7)DocsLinux manual pages(opens in a new tab)
- Control Group v2DocsLinux kernelcpu.max, cpu.weight, memory.max, io.max, cgroup.freeze(opens in a new tab)
- Freezer SubsystemDocsLinux kernel(opens in a new tab)
- Resource constraintsDocsDocker(opens in a new tab)
- Docker Engine securityDocsDocker(opens in a new tab)
- Setting limits for applicationsDocsRed Hat (RHEL 8)(opens in a new tab)
- CRIU: Checkpoint/Restore In UserspaceDocsCRIU project(opens in a new tab)
- Firecracker: Lightweight Virtualization for Serverless ApplicationsPaperUSENIX NSDI 2020 (Agache et al.)(opens in a new tab)
- Performance Overhead Comparison between Hypervisor and Container based VirtualizationPaperIEEE AINA 2017 (Li, Kihl, Lu, Andersson)(opens in a new tab)
- An updated performance comparison of virtual machines and Linux containersPaperIEEE ISPASS 2015 (Felter et al.)Further reading on VM and container performance(opens in a new tab)