Majid Al-RaimiHypervisors, containers, namespaces and cgroups

COE 558Lecture 02Part 04

Hypervisors, containers, namespaces and cgroups

Compares Type 1 and Type 2 hypervisors, then explains OS-level virtualization (containers) built from Linux namespaces and control groups.

Concepts
6
Slides
22-28
Reading
36 min
Understood
0/6 concepts

Why this part matters

Every cloud service in the rest of this course, from EC2 and Lambda to Kubernetes on EKS or GKE and edge nodes running k3s, is built from one of two isolation designs. Either a hypervisor splits the hardware into virtual machines, or one shared kernel is split into containers by namespaces and cgroups.

The choice between them decides overhead, density, start-up time and how strong the security boundary is. Comparing Type 1, Type 2 and containers, and telling namespaces from cgroups, are the likely exam questions. The same choice comes back when you size edge devices for the research project, where a small gateway may not have the memory to boot a full guest OS for every service.

By the end you can

  1. Place Type 1 and Type 2 hypervisors in the software stack and name real examples of each, including where KVM and Hyper-V fit.
  2. Compare Type 1 and Type 2 on overhead, resource management, administration and isolation, and explain why cloud providers choose Type 1.
  3. Explain OS-level virtualization: what containers share, what they isolate, and why their security boundary is weaker than a VM's.
  4. Distinguish namespaces (what a process sees) from cgroups (what a process uses), and name the main namespace types and cgroup functions.
  5. Read a cgroup CPU quota and weight and predict how a container behaves under contention and at its memory limit.

Where the hypervisor sits: bare metal or hosted

Start with a laptop running VirtualBox with Ubuntu inside it. When you press the power button, macOS boots first. VirtualBox is then just another application, next to Safari, and the whole Ubuntu guest lives inside that application's process. Now compare a server in an AWS data center that hosts EC2 instances. Nothing general purpose boots before the hypervisor. It is the first thing on the machine, and it hands out CPU and memory to the guests directly.

Both machines do Virtualization, and both run a Virtual machine monitor (VMM) that makes each Virtual machine (VM) believe it has a real computer. The difference is where the VMM sits, and that comes down to one question: who owns the hardware?

Two placements

  • A Type 1 hypervisor (bare metal, or native) is the lowest software layer on the machine. It runs directly on the hardware and acts as a specialized OS whose only job is to run guests. Red Hat describes it as a hypervisor that "runs directly on the host's hardware to manage guest operating systems".
  • A Type 2 hypervisor (hosted) runs on a conventional OS "as a software layer or application". The host OS owns the hardware, and the VMM is one of its processes, with less privilege than the OS itself. VirtualBox's own manual calls it "a so-called hosted hypervisor, sometimes referred to as a type 2 hypervisor".
Type 1 (bare metal)
Hardware, then hypervisor, then guests

Guest OS 1 to N run directly on the hypervisor, which owns CPU, memory, storage and network.

add a host OS
Type 2 (hosted)
Hardware, then host OS, then VMM, then guests

The VMM is an app on the host OS. Each guest request crosses the VMM and the host kernel.

Type 1 owns the hardware; Type 2 asks a host OS for it
What stands between the workload and the hardware: a hypervisor, a host OS plus a VMM, or one shared kernel

Real examples, and where the labels get blurry

VMware's Type 1 hypervisor is ESXi, described as "the hypervisor [that] runs virtual machines"; vCenter Server is the central management service for many ESXi hosts, and vSphere is the name of the whole suite. On the hosted side sit VirtualBox, VMware Workstation and Parallels Desktop.

Picture a lab PC running VirtualBox for a coursework VM, next to an EC2 c7i instance. When the coursework VM compiles code, its virtual CPU is just a thread that competes with Chrome and the antivirus for time on the host scheduler. The EC2 instance gets CPUs and memory partitioned for it by a minimal hypervisor, with no desktop OS in the way.

That picture gives the whole comparison. Because a Type 1 hypervisor owns the hardware, it adds less overhead (the slide says about 5%), manages resources directly and isolates guests more strongly, but it needs more skill to install and operate. A Type 2 hypervisor reaches the hardware only through the host OS, so it costs more (about 10%), controls resources indirectly and isolates less, but anyone can install it.

PropertyType 1 (bare metal)Type 2 (hosted)
Runs onDirectly on the hardware, as the lowest software layerAs a process of a general purpose host OS
Approx. overhead≈5%≈10%
Resource managementDirect: the hypervisor schedules CPUs and partitions memory itselfIndirect: every request goes through the host OS scheduler and drivers
AdministrationNeeds more skill: a dedicated server, remote management, clustersEasy: install it like any desktop app
IsolationStronger: only a small hypervisor sits between tenantsWeaker: a compromise of the host OS exposes every guest
Typical usersCloud providers and enterprise data centersDevelopers, students, testers on laptops
ExamplesVMware ESXi, Microsoft Hyper-V, KVM (debated), AWS NitroOracle VirtualBox, VMware Workstation, Parallels Desktop
Type 1 and Type 2 hypervisors side by side

Why cloud providers pick Type 1

A Cloud service provider (CSP) sells slices of the same physical server to strangers. Every percent of overhead is capacity it cannot sell, and every extra layer between tenants is attack surface. AWS shows where this leads. Its Nitro Hypervisor is "a custom-developed, minimized hypervisor based on KVM", cut down on purpose: it has "no networking stack, no general-purpose file system implementations, and no peripheral device driver support". Networking and storage are offloaded to dedicated Nitro Cards and exposed to guests through SR-IOV. The hypervisor does little more than partition CPU and memory, which is the Type 1 idea taken to its limit, and the design is lean enough that AWS can also sell bare metal instances.

On one Ubuntu server, run docker run nginx and then docker run mongo. Each starts in well under a second, because neither boots an operating system. Both are ordinary Linux processes making system calls into the same host kernel. Run uname -r inside either one and it prints the host's kernel version, because there is no other kernel to ask.

That is OS-level virtualization. There is no hypervisor and no guest kernel. The line that matters is the System call interface, the entry point through which user-space code asks the kernel for anything. Above it, each environment has its own isolated user space: the application and its libraries. Below it, everything is shared: the kernel, and the hardware (or the virtual hardware, if the host is itself a VM). These isolated environments are called a Container.

The design has three consequences. Containers are light and fast, because starting one is starting a process. They are tied to one OS family, because the shared kernel must understand their system calls. NIST is blunt: "a Linux host can only run containers built for Linux". And their boundary is weaker than a VM's, because a shared kernel "invariably results in a larger inter-object attack surface than seen with hypervisors". One kernel bug can be a door between every container on the host.

PropertyVirtual machineContainer
Kernel per instanceYes: each VM boots its own guest kernelNo: every container calls into the one host kernel
Start-upSeconds to minutes (an OS boot)Milliseconds to about a second (a process start)
Isolation boundaryHypervisor and virtual hardware: small, concreteSystem call interface of a shared kernel: large attack surface
OS familyAny OS the virtual hardware supportsSame family as the host kernel only
Image sizeGigabytes (full OS)Megabytes (application plus libraries)
Virtual machines and containers

Two problems, two kernel features

Take the slide's own example: one server runs a database and a web server as plain processes. Two things can go wrong. First, the web server can list the database's processes, send them signals, read its files and bind its ports, so whoever compromises one has a view of the other. Second, a heavy query makes the database use every CPU core, and the website stalls.

These are different problems, and Linux fixes each with a different kernel feature. The first is about visibility. A Namespace gives a group of processes its own view of a kernel resource, so the web server cannot even see the database's processes, network stack or mount points. The second is about consumption. Control groups (cgroups) put quotas, weights and accounting on how much CPU, memory and I/O a group can use. Docker's security guide gives the same split: namespaces "provide the first and most straightforward form of isolation", while cgroups "implement resource accounting and limiting".

ProblemFixKernel featureQuestion it answers
Processes see and touch each other (PIDs, ports, files)Give each group its own view of kernel resourcesNamespacesWhat can I see?
Processes compete for the same CPU, memory and I/OPut quotas, weights and accounting on resource useControl groups (cgroups)How much can I use?
One problem per kernel feature
Namespaces draw the walls (what each process sees); cgroups set the cap (what each process uses)

Neither feature is new. The first namespace (mount) arrived in Linux 2.4.19 in 2002, most of the rest landed between 2.6.15 and 2.6.26 (July 2008), and user namespaces were only completed in 3.8 (2013), which made unprivileged containers practical. Cgroups were started in 2006 and merged in 2.6.24. A Container is simply these two combined, plus a root filesystem image. NIST describes the second half with a concrete case: a host with 10 GB of memory can give 1 GB to each of nine containers, so no single one can starve the rest.

Namespaces: what a process can see

You can make a Namespace by hand on any Linux machine. The command below starts a shell in a new PID namespace and mounts a fresh /proc for it. Inside, ps a shows only two processes, and the shell believes it is PID 1.

sudo unshare --fork --pid --mount-proc bash
ps a
Start a shell in a new PID namespace
PID TTY      STAT   TIME COMMAND
      1 pts/0    S      0:00 bash
      8 pts/0    R+     0:00 ps a
Output of ps a inside the namespace

From a second terminal on the host, the same bash appears in ps -A under a large global PID such as 24517. Nothing was copied and no second kernel was started. The kernel simply shows that process a different numbering. NIST gives the container version: ps -A inside an Apache Container lists only httpd.

One process list, two views: the host sees global PIDs, the namespace sees its own numbering from 1

The rule

The Linux manual defines a namespace as something that "wraps a global system resource in an abstraction that makes it appear to the processes within the namespace that they have their own isolated instance". There is one kernel and many views. Every process belongs to exactly one namespace of each type, and a container runtime such as Docker creates a fresh set (all but user and time, by default) for every container. You can see a process's memberships as symbolic links in /proc/<pid>/ns/ (since Linux 3.8), and a tool can join an existing namespace with setns(2), which is how docker exec gets a shell into a running container.

NamespaceFlagWhat it isolates
CgroupCLONE_NEWCGROUPThe cgroup root directory a process sees
IPCCLONE_NEWIPCSystem V IPC objects and POSIX message queues
NetworkCLONE_NEWNETNetwork devices, IP stacks, routing tables, ports
MountCLONE_NEWNSMount points, so each container has its own filesystem tree
PIDCLONE_NEWPIDProcess ID numbers, starting again at 1
TimeCLONE_NEWTIMEBoot and monotonic clock offsets
UserCLONE_NEWUSERUser and group IDs, so root inside can be unprivileged outside
UTSCLONE_NEWUTSHostname and NIS domain name
The eight Linux namespace types (namespaces(7))

The network namespace is what gives each container "its own networking stack, including unique interfaces and IP addresses", in NIST's words, so two containers can both listen on port 80. The user namespace, when enabled, lets a process be root inside its container while mapping to an unprivileged user on the host.

Control groups: what a process can use

On a host with 2 CPUs, start MongoDB with docker run --cpus=1.5 --memory=512m mongo. Docker writes two numbers into the container's cgroup: a CPU quota of 150000 µs per 100000 µs period, and a memory ceiling of 512 MiB. If MongoDB grows past that ceiling and the kernel cannot reclaim memory, the OOM killer kills a process inside that cgroup only. The web server next to it never notices.

docker run -d --name db --cpus=1.5 --memory=512m mongo
cat /sys/fs/cgroup/system.slice/docker-$(docker inspect -f '{{.Id}}' db).scope/cpu.max
cat /sys/fs/cgroup/system.slice/docker-$(docker inspect -f '{{.Id}}' db).scope/memory.max
Set limits, then read them back from cgroup v2
150000 100000
536870912
cpu.max and memory.max (512 MiB in bytes)

How cgroups work

The manual says cgroups "allow processes to be organized into hierarchical groups whose usage of various types of resources can then be limited and monitored". Each kind of resource is enforced by a controller, which Red Hat defines as "a kernel subsystem that represents a single resource, such as CPU time, memory, network bandwidth or disk I/O". The main controllers are cpu, memory, io, pids and cpuset; freezing, a separate freezer controller in cgroup v1, is the core file cgroup.freeze in v2. Cgroup v2, official since Linux 4.5, puts everything in one tree mounted at /sys/fs/cgroup, where every setting is a small file you can read and write.

The slide lists four functions, and each maps to concrete files:

The four cgroup functions and the cgroup v2 files behind them

Limiting
Hard caps: cpu.max, memory.max, io.max, pids.max
Prioritization
Shares under contention: cpu.weight, io.weight
Accounting
Measured use for monitoring and billing: cpu.stat, memory.current, io.stat
Control
Freezing and thawing a whole group: cgroup.freeze (the basis for checkpoint and restart with CRIU)

Worked example

Reading a CPU quota and a CPU weight

  1. Turn cpu.max into CPUs

    The file holds 150000 100000: a quota of CPU time, then the period it applies to. The format is $MAX $PERIOD, and the default is max 100000 (no limit).

    CPUs=quotaperiod=150000 μs100000 μs=1.5\begin{aligned} \text{CPUs} &= \frac{\text{quota}}{\text{period}} \\ &= \frac{150000\,\mu\text{s}}{100000\,\mu\text{s}} = 1.5 \end{aligned}
  2. Compare with the machine

    On a 2-CPU host that is 1.5 / 2 = 75% of all CPU time. It is a hard ceiling: once the group has used 150 ms of CPU time in a 100 ms period, it waits for the next period, even if the machine is otherwise idle.

  3. Add prioritization

    Now two busy containers have --cpu-shares 1024 and 512 (cgroup v2 calls this cpu.weight, default 100, range 1 to 10000). When they compete, each gets its weight divided by the sum of weights:

    si=wi∑jwj10241536≈23,5121536≈13\begin{gathered} s_i = \frac{w_i}{\sum_j w_j} \\ \frac{1024}{1536} \approx \tfrac{2}{3}, \quad \frac{512}{1536} \approx \tfrac{1}{3} \end{gathered}

    This is a soft limit. If the second container goes idle, the first may use all the CPU it can get.

  4. Result

    Limits are ceilings that hold even on an idle machine; weights are shares that matter only under contention. That is how to read the slide's diagram where four groups each get 25%: with equal weights, each gets a quarter only while all four are busy.
A quota is time per period: 150 ms of CPU in every 100 ms, spread across cores, then the group waits

Recall

In one sentence each, where does a Type 1 and a Type 2 hypervisor run, and who owns the hardware?

Type 1 runs directly on the hardware and owns it (it acts as the OS). Type 2 runs as a process on a host OS, which owns the hardware, so the VMM gets resources only indirectly and at lower privilege.

Recall

Why do cloud providers choose Type 1 hypervisors?

Lower overhead (the slide says about 5% versus 10%), direct resource management and stronger isolation between tenants. Example: the AWS Nitro Hypervisor, a minimal KVM-based Type 1.

Recall

What is shared and what is isolated in OS-level virtualization?

The kernel (below the system call interface) and the hardware are shared. Each container's application and libraries are isolated. There is no hypervisor and no guest kernel.

Recall

Namespaces and cgroups: which question does each answer?

Namespaces answer "what can this process see" (PIDs, network, mounts, hostname, users, IPC, time, cgroup root). Cgroups answer "how much can it use" (CPU, memory, I/O, PIDs), through limiting, prioritization, accounting and control.

Recall

A container gets --cpus=1.5. What does cpu.max hold, and what does it mean?

150000 100000: up to 150 ms of CPU time in every 100 ms period, which is 1.5 CPUs, a hard ceiling.

Recall

Two busy containers have cpu.weight 200 and 100 and no cpu.max. What share does each get, and what happens when the second goes idle?

About two thirds and one third while both are busy. Weights matter only under contention, so once the second goes idle the first may use all the CPU.

Quick check

A cloud provider wants minimal overhead and strong isolation between tenant VMs on each server. Which virtualization layer should it deploy?

Quick check

A database and a web server share one Linux host. The database starts using every CPU core and the website slows down. Which kernel feature fixes this?

Quick check

Inside a container, ps shows only two processes and your app has PID 1. What produces this view?

Quick check

Why can a Linux host run a Windows VM but not a native Windows container?

Recap

If you remember nothing else

  • Type 1 hypervisors run on bare metal and act as the OS (ESXi, Hyper-V, KVM). Type 2 run as an app on a host OS (VirtualBox, VMware Workstation, Parallels).
  • Type 1: about 5% overhead, direct resource control, stronger isolation, harder to run. Type 2: about 10% overhead, indirect control, easier, less isolation.
  • Cloud providers use Type 1. AWS Nitro is a minimal KVM-based Type 1 hypervisor that offloads I/O to dedicated cards.
  • Containers use no hypervisor. They share the host kernel below the system call interface and isolate only their user space.
  • Namespaces control what a process can see: PID, network, mount, UTS, IPC, user, cgroup and time.
  • Cgroups control what a process can use: limiting, prioritization, accounting and control (freezing a whole group, which CRIU builds on to checkpoint and restart).
  • A container is several namespaces plus cgroups plus a root filesystem. In the cloud, containers usually run inside VMs.

Sources