CHAPTER 07 40 MIN READ ADVANCED

Container and Docker Security: Fundamentals, Escape Techniques & Hardening

Containers share the host kernel instead of virtualizing hardware, which is what makes them lighter than a virtual machine. That same design means a container boundary is a different, and sometimes weaker, kind of boundary than most people assume when they hear the word "isolation."

container fundamentals Docker security container escape Kubernetes basics

Container Fundamentals: Namespaces and cgroups

Everything a container is rests on two Linux kernel mechanisms, not on some separate container technology. Namespaces control what a process can see. cgroups control what a process can use.

  • Namespaces: isolate a process's view of the system, its own process list (PID namespace), its own network stack (network namespace), its own mount points (mount namespace), and more. A process inside a namespace generally cannot see or touch what is outside it.
  • cgroups (control groups): limit and account for the resources a process or group of processes can consume, CPU time, memory, disk I/O. cgroups are why one noisy container can be stopped from starving every other workload on the same host.

A container, in practice, is just a regular Linux process launched with a specific set of namespaces and cgroup limits applied to it. There is no separate "container kernel" running underneath.

AspectContainersVirtual Machines
KernelShare the host's single kernelEach VM runs its own separate guest kernel
Isolation mechanismNamespaces and cgroups, enforced in the shared kernelA hypervisor virtualizing hardware, with a stronger boundary between guests
Startup weightLightweight, starts in secondsHeavier, boots a full guest OS
Blast radius of a kernel flawCan potentially affect every container sharing that kernelA guest kernel flaw is contained to that one VM, not the host or its siblings
The practical consequence: a kernel-level vulnerability is a shared-fate risk for every container on that host in a way it simply is not for separate VMs, because VMs do not share a kernel with each other or with the hypervisor's own workloads.

Docker Architecture and Risk

The Docker daemon (dockerd) is the process that actually creates and manages containers, and it typically runs as root. Anything that can talk to that daemon is, functionally, talking to something with full host privileges.

ConfigurationRisk
Privileged container (--privileged)Disables most of the isolation containers normally provide, giving the container far broader access to the host than a default container would have
Host network or PID namespace sharingThe container can see host-level network traffic or host processes it would not normally be able to see, weakening the isolation those namespaces exist to provide
Mounting the Docker socket (/var/run/docker.sock) into a containerA well-known technique that gives the container the ability to control the host's Docker daemon directly, which is effectively equivalent to host root

None of these three configurations are exotic. They show up in real deployments, often because a monitoring tool, a CI runner, or a debugging container needed broad access and nobody scoped it back down afterward.

Container Escape Technique Categories

A container escape is any technique that lets a process break out of its container's isolation and reach the host or other containers. These fall into a small number of conceptual categories, described here at the category level rather than as working exploit steps.

CategoryWhat It Relies On
Privileged-container abuseThe relaxed isolation of a --privileged container, which removes many of the normal restrictions between the container and the host
Sensitive mount abuseA mounted host path, or the Docker socket, that gives the container direct access to host resources it should not be able to reach
Kernel vulnerability exploitationA flaw in the shared kernel itself, unrelated to how the container was configured, and directly connected to the shared-kernel risk described in the fundamentals section above
Why the category matters more than the exact technique: the first two categories are fixable through configuration choices covered later in this chapter. The third is a kernel patching and update problem, not a container configuration problem.

Kubernetes Security Basics

Kubernetes is the layer that schedules and manages containers across a fleet of hosts, and it introduces its own security concepts on top of the container-level risks already covered.

  • Pods: the smallest deployable unit in Kubernetes, one or more containers scheduled together on the same node, sharing some resources between them.
  • RBAC (role-based access control): governs what a given identity is allowed to do against the Kubernetes API, including what a pod's own service account is permitted to do.
Token scope matters: a compromised pod's service account token can sometimes be used to interact with the Kubernetes API directly. An overly broad RBAC role attached to that service account turns a single compromised pod into a foothold against the whole cluster's API.
This chapter covers Kubernetes concepts only, not hunting methodology. See Threat Hunting, Chapter 11: Container and Kubernetes Hunting for the ATT&CK-for-Containers hunting methodology, Kubernetes audit log analysis, and Falco-based runtime detection.

Container Hardening Basics

Most of the hardening guidance here directly closes off the risks and escape categories already covered in this chapter.

PracticeWhy
Avoid privileged containers unless there is a specific, understood reasonMinimizes the escape surface described in the escape technique categories section
Never mount the Docker socket into a container that doesn't need to control Docker itselfRemoves the most direct path from container access to effective host root
Run processes inside containers as a non-root userLimits what an attacker can do even if they do get code running inside the container
Use read-only root filesystems where the application allows itPrevents an attacker from writing new tools or persistence into the container itself
Keep base images minimal and patchedFewer installed tools means less for an attacker to abuse if they get in, the same living-off-the-land theme as shell tradecraft
Default to least privilege: every one of these practices is a variation on the same idea, give a container exactly the access it needs to do its job, and nothing more. That single principle closes most of the risk surface covered in this chapter.

Key Takeaways

  • Containers are built on two kernel mechanisms: namespaces isolate what a process can see, cgroups limit and account for what a process can use.
  • Containers share the host's single kernel, unlike VMs, so a kernel vulnerability can potentially affect every container on that host.
  • The Docker daemon typically runs as root, and mounting the Docker socket into a container is effectively equivalent to giving that container host root.
  • Container escapes fall into three categories: privileged-container abuse, sensitive mount abuse, and kernel vulnerability exploitation.
  • Kubernetes adds pods and RBAC on top of container-level risk, and a compromised pod's service account token is only as dangerous as the RBAC role attached to it.
  • Hardening comes down to least privilege: avoid privileged mode, never expose the Docker socket unnecessarily, run as non-root, use read-only filesystems, and keep images minimal.

Knowledge Check

Click an answer to reveal the explanation.

Mounting /var/run/docker.sock into a container effectively grants that container what?

Correct answer: C. The Docker daemon runs as root, so anything able to issue commands to it through the socket can direct it to do anything root-level Docker access allows on the host.

What is the core difference between what namespaces do and what cgroups do?

Correct answer: B. Namespaces control visibility, such as a container's own process list or network stack. cgroups control and account for consumption of resources like CPU and memory.

Why is a shared-kernel vulnerability a container-specific risk that virtual machines don't share in the same way?

Correct answer: B. Every container on a host is a process running against that same shared kernel. VMs each have their own guest kernel, so a kernel flaw in one VM does not extend to its siblings or the host the way it can across containers.