Container and Docker Security: Fundamentals, Escape Techniques & Hardening
Containers share the host kernel instead of virtualizing hardware, which is what makes them lighter than a virtual machine. That same design means a container boundary is a different, and sometimes weaker, kind of boundary than most people assume when they hear the word "isolation."
Container Fundamentals: Namespaces and cgroups
Everything a container is rests on two Linux kernel mechanisms, not on some separate container technology. Namespaces control what a process can see. cgroups control what a process can use.
- Namespaces: isolate a process's view of the system, its own process list (PID namespace), its own network stack (network namespace), its own mount points (mount namespace), and more. A process inside a namespace generally cannot see or touch what is outside it.
- cgroups (control groups): limit and account for the resources a process or group of processes can consume, CPU time, memory, disk I/O. cgroups are why one noisy container can be stopped from starving every other workload on the same host.
A container, in practice, is just a regular Linux process launched with a specific set of namespaces and cgroup limits applied to it. There is no separate "container kernel" running underneath.
| Aspect | Containers | Virtual Machines |
|---|---|---|
| Kernel | Share the host's single kernel | Each VM runs its own separate guest kernel |
| Isolation mechanism | Namespaces and cgroups, enforced in the shared kernel | A hypervisor virtualizing hardware, with a stronger boundary between guests |
| Startup weight | Lightweight, starts in seconds | Heavier, boots a full guest OS |
| Blast radius of a kernel flaw | Can potentially affect every container sharing that kernel | A guest kernel flaw is contained to that one VM, not the host or its siblings |
Docker Architecture and Risk
The Docker daemon (dockerd) is the process that actually creates and manages containers, and it typically runs as root. Anything that can talk to that daemon is, functionally, talking to something with full host privileges.
| Configuration | Risk |
|---|---|
Privileged container (--privileged) | Disables most of the isolation containers normally provide, giving the container far broader access to the host than a default container would have |
| Host network or PID namespace sharing | The container can see host-level network traffic or host processes it would not normally be able to see, weakening the isolation those namespaces exist to provide |
Mounting the Docker socket (/var/run/docker.sock) into a container | A well-known technique that gives the container the ability to control the host's Docker daemon directly, which is effectively equivalent to host root |
None of these three configurations are exotic. They show up in real deployments, often because a monitoring tool, a CI runner, or a debugging container needed broad access and nobody scoped it back down afterward.
Container Escape Technique Categories
A container escape is any technique that lets a process break out of its container's isolation and reach the host or other containers. These fall into a small number of conceptual categories, described here at the category level rather than as working exploit steps.
| Category | What It Relies On |
|---|---|
| Privileged-container abuse | The relaxed isolation of a --privileged container, which removes many of the normal restrictions between the container and the host |
| Sensitive mount abuse | A mounted host path, or the Docker socket, that gives the container direct access to host resources it should not be able to reach |
| Kernel vulnerability exploitation | A flaw in the shared kernel itself, unrelated to how the container was configured, and directly connected to the shared-kernel risk described in the fundamentals section above |
Kubernetes Security Basics
Kubernetes is the layer that schedules and manages containers across a fleet of hosts, and it introduces its own security concepts on top of the container-level risks already covered.
- Pods: the smallest deployable unit in Kubernetes, one or more containers scheduled together on the same node, sharing some resources between them.
- RBAC (role-based access control): governs what a given identity is allowed to do against the Kubernetes API, including what a pod's own service account is permitted to do.
Container Hardening Basics
Most of the hardening guidance here directly closes off the risks and escape categories already covered in this chapter.
| Practice | Why |
|---|---|
| Avoid privileged containers unless there is a specific, understood reason | Minimizes the escape surface described in the escape technique categories section |
| Never mount the Docker socket into a container that doesn't need to control Docker itself | Removes the most direct path from container access to effective host root |
| Run processes inside containers as a non-root user | Limits what an attacker can do even if they do get code running inside the container |
| Use read-only root filesystems where the application allows it | Prevents an attacker from writing new tools or persistence into the container itself |
| Keep base images minimal and patched | Fewer installed tools means less for an attacker to abuse if they get in, the same living-off-the-land theme as shell tradecraft |
Key Takeaways
- Containers are built on two kernel mechanisms: namespaces isolate what a process can see, cgroups limit and account for what a process can use.
- Containers share the host's single kernel, unlike VMs, so a kernel vulnerability can potentially affect every container on that host.
- The Docker daemon typically runs as root, and mounting the Docker socket into a container is effectively equivalent to giving that container host root.
- Container escapes fall into three categories: privileged-container abuse, sensitive mount abuse, and kernel vulnerability exploitation.
- Kubernetes adds pods and RBAC on top of container-level risk, and a compromised pod's service account token is only as dangerous as the RBAC role attached to it.
- Hardening comes down to least privilege: avoid privileged mode, never expose the Docker socket unnecessarily, run as non-root, use read-only filesystems, and keep images minimal.
Knowledge Check
Click an answer to reveal the explanation.
Mounting /var/run/docker.sock into a container effectively grants that container what?
What is the core difference between what namespaces do and what cgroups do?
Why is a shared-kernel vulnerability a container-specific risk that virtual machines don't share in the same way?