CHAPTER 11 40 MIN READ ADVANCED

Container and Kubernetes Hunting

A container is not a lightweight virtual machine. It is a set of Linux namespaces and cgroups sharing one kernel with every other container on the host, orchestrated by a control plane that itself is reachable over the network as an API. That combination creates attack surface no endpoint-hunting chapter in this course has touched: escaping a container means reaching the host kernel directly, and compromising one over-privileged pod's service account can mean compromising the entire cluster through the Kubernetes API. This chapter builds the hunting model for that environment.

Kubernetes Falco ATT&CK Containers

Why Containers Need a Different Hunting Model

Every technique covered in Chapters 1 through 9 assumes a single operating system instance with its own kernel, its own process tree, and Sysmon or an EDR agent watching it directly. Kubernetes breaks that model in three specific ways that change what a hunter should look for and where.

Three Ways Kubernetes Breaks the Model

  • Containers share a kernel. Unlike a hypervisor-based VM, a container has no hardware-enforced isolation boundary. A kernel vulnerability, a misconfigured capability, or a mounted host resource can let a process inside a container directly affect the host or every other container running on it. This is the entire premise of T1611, Escape to Host.
  • The orchestrator is itself an API-reachable attack surface. The Kubernetes API server is a control plane that can create, delete, and exec into workloads across the entire cluster. A credential that grants API access (a stolen kubeconfig, an over-permissioned service account token) is functionally equivalent to domain admin in a Windows environment, except the "domain" is every workload the cluster runs.
  • Workloads are ephemeral by design. A compromised pod might live for minutes before the orchestrator reschedules it. Traditional incident response assumptions (the compromised host is still there tomorrow for forensic imaging) do not hold. Hunting has to work from the audit trail and runtime telemetry, because the pod itself may already be gone.

Applying the Same ATT&CK Discipline

This chapter maps to the same ATT&CK-driven hypothesis discipline from Chapter 2, using the ATT&CK for Containers matrix published by the Center for Threat-Informed Defense, plus the container-relevant sub-techniques already present in TTPHUNT: T1053.007 (Container Orchestration Job) and T1059.013 (Container CLI/API). Where this chapter goes further is the cluster-level control plane: RBAC, the API server audit log, and container runtime behavior.

Note: This chapter uses Kubernetes as the worked example because it is the dominant container orchestrator in enterprise environments. The underlying concepts (shared-kernel escape, orchestrator-as-attack-surface, ephemeral workloads) apply equally to Docker Swarm, Nomad, and ECS, with different API and audit log specifics.

ATT&CK for Containers Overview

The Center for Threat-Informed Defense published ATT&CK for Containers as an extension to the Enterprise matrix, adding techniques specific to container and orchestration abuse rather than forcing container attacks into host-centric technique definitions that did not quite fit.

Five Anchor Techniques

Technique What It Covers Primary Hunt Surface
T1610: Deploy Container Adversary deploys a new container (often from a malicious or unmodified public image) to execute code, establish persistence, or evade defenses that were watching the existing workloads Kubernetes audit log: create verb on pods/deployments from an unexpected principal or CI/CD identity outside its normal deploy window
T1611: Escape to Host Adversary breaks the container boundary to execute code on the underlying node, gaining access to every other container and the host's own credentials Privileged pod specs, mounted hostPath/Docker socket, runtime (Falco) syscall detections for namespace/capability abuse
T1612: Build Image on Host Adversary sends a build request directly to the container runtime's API (e.g., the Docker daemon socket) to build and run an image locally, bypassing defenses that only scan images pulled from a registry Docker/containerd API build-request logs; runtime detection for unexpected daemon socket access
T1613: Container and Resource Discovery Adversary enumerates containers, pods, namespaces, and cluster resources to understand what they have landed in and what else is reachable Audit log list/get verbs against pods, namespaces, secrets at unusual volume from a single identity
T1552.007: Unsecured Credentials: Container API Adversary queries the Docker API or Kubernetes API (frequently via an over-permissioned pod service account token) to retrieve credentials, secrets, or logs containing sensitive data Audit log access to the secrets resource; service account token usage from a pod that has no legitimate reason to read secrets outside its own namespace
Tip: T1609, Container Administration Command, covers the simplest and most common technique in practice: an adversary with API or kubeconfig access simply runs kubectl exec into a running pod, or uses the Docker CLI/API directly, to execute commands. It requires no exploit and no misconfiguration beyond overly broad access, which is exactly why pods/exec auditing (Section 3) is one of the highest-value hunts in this chapter.

Kubernetes Audit Log Hunting

The Audit Log as Primary Data Source

The Kubernetes API server can emit a structured audit log (audit.k8s.io/v1) for every request it processes, if an audit policy is configured. This is the single most important data source in this chapter: every kubectl command, every controller reconciliation, and every service account API call passes through the API server and can be captured here.

Core fields to build hunt queries around: verb (get, list, create, update, patch, delete, deletecollection), objectRef (the resource type, namespace, and name being acted on, including a subresource field for things like exec and portforward), user (the identity making the call, a human via SSO, or a service account), and sourceIPs.

Example Audit Log Entry

JSON
{
  "apiVersion": "audit.k8s.io/v1",
  "kind": "Event",
  "level": "Metadata",
  "verb": "create",
  "objectRef": {
    "resource": "pods",
    "namespace": "production",
    "name": "web-frontend-7c9f8d-xk2lp",
    "apiVersion": "v1",
    "subresource": "exec"
  },
  "user": { "username": "jane.doe@corp.com", "groups": ["system:authenticated"] },
  "sourceIPs": ["203.0.113.45"],
  "requestURI": "/api/v1/namespaces/production/pods/web-frontend-7c9f8d-xk2lp/exec?command=/bin/sh"
}

Three Hunt Patterns

These cover most of the value in this log source:

  • Exec into production pods. Filter for objectRef.subresource: "exec" or "attach" against namespaces that should never need interactive debugging (production, especially anything handling regulated data). A legitimate on-call engineer execs in occasionally; a pattern of repeated execs across many pods from one identity in a short window looks like discovery or lateral movement, not debugging.
  • Anonymous or unauthenticated API access. A user.username of system:anonymous succeeding against any resource indicates the API server (or an exposed component like the kubelet's read-only port) is reachable without authentication, a critical exposure regardless of what specific resource was accessed.
  • RBAC self-escalation. Watch for create/update verbs against clusterrolebindings, rolebindings, or clusterroles performed by an identity that is not a documented cluster administrator. Binding oneself (or a controlled service account) to cluster-admin is the single most direct privilege escalation path in Kubernetes.
Warning: An empty or near-empty audit log is not evidence of a quiet cluster. Kubernetes audit logging is opt-in and, at the default Metadata level, may exclude the request/response bodies you need for some investigations. Before trusting a hunt result of "no matches," verify the audit policy actually covers the verbs, resources, and namespaces in scope: the same Equip-stage discipline from Chapter 3's SANS PAM coverage.

Runtime Detection with Falco

Why Falco Fills the Gap

The Kubernetes audit log tells you what was requested through the API. It has no visibility into what a process does once it is already running inside a container: a reverse shell spawned by an exploited application, a write to a sensitive host-mounted path, or a process reading the pod's own service account token to exfiltrate it. That gap is covered by syscall-level runtime security tooling, of which Falco (a CNCF graduated project) is the de facto standard.

Falco rules match on syscall events using a condition expression built from fields like proc.name, container.id, fd.name, and evt.type, reusable via macros and lists. The canonical example, "Terminal shell in container," illustrates the pattern every container runtime hunt should understand:

Example Falco Rule

FALCO RULE
- rule: Terminal shell in container
  desc: Detect a shell being spawned in a container
  condition: > spawned_process and container and shell_procs and proc.tty != 0
  output: > Shell spawned in container (user=%user.name
    container_id=%container.id container_name=%container.name
    shell=%proc.name parent=%proc.pname cmdline=%proc.cmdline)
  priority: WARNING

This single rule is disproportionately valuable because most production container images are built without an interactive shell need in mind: a shell spawning inside a minimal application container, attached to a TTY, is rarely a legitimate operational action and very frequently the first sign of a working remote code execution exploit.

More Falco Hunt Patterns

Falco Hunt Pattern What It Catches
Shell spawned in a container with no interactive-shell business need Post-exploitation shell from a web application RCE, T1609 Container Administration Command abuse
Write below /etc, sensitive mount paths, or a read-only root filesystem attempt Persistence via config tampering; also flags containers that were supposed to run with a read-only root filesystem but do not
Outbound connection to an unexpected port/destination from a container that should have no egress need C2 callback, data exfiltration from a workload with a narrow expected network profile
Process attempting to read the container's own service account token file outside expected client libraries Credential harvesting consistent with T1552.007, Unsecured Credentials: Container API
Tip: Falco ships with a substantial default rule set covering common abuse patterns out of the box. The highest-value hunting work is not writing rules from scratch, but tuning the default set against your own baseline: which containers legitimately need shell access for debugging sidecars, which need broad egress, and excluding those explicitly rather than silencing the rule tenant-wide.

Container Escape and Privileged Pod Hunting

Why Pod Specs Matter More Than Exploits

Most real-world container escapes (T1611) do not require a kernel zero-day. They require a pod spec that was granted more access than the workload needed, deployed by a team optimizing for "it works" over least privilege. Hunting for escape risk is largely a matter of hunting the pod specification itself, before any exploit ever runs.

Escape-Enabling Pod Spec Settings

Pod Spec Setting Why It Enables Escape
securityContext.privileged: true Disables nearly all container isolation, granting access to all host devices and effectively equivalent to root on the host
hostPath volume mount, especially of /, /etc, or /var/run/docker.sock Gives the container direct read/write access to the host filesystem or, in the case of the Docker socket, full control over the container runtime itself (equivalent to root on the host)
hostNetwork: true Removes network namespace isolation, exposing the container to every network interface and service the host itself can reach, and vice versa
hostPID: true Exposes the host's entire process namespace to the container, allowing it to see and potentially signal/ptrace processes running outside its own namespace
Added Linux capabilities (SYS_ADMIN, SYS_PTRACE, SYS_MODULE) Each grants a specific piece of root-equivalent kernel functionality; SYS_ADMIN in particular exposes a large surface of historically exploitable escape primitives

Hunting the Spec, Not the Log

Because these are static configuration facts rather than runtime events, the hunt is a periodic sweep, not a log query: enumerate every pod spec in the cluster (via kubectl get pods -o json across all namespaces, or an admission-controller policy engine's audit mode) and flag any workload carrying one of the settings above without a documented, reviewed exception.

The Docker socket mount is worth tracing end to end, since it is one of the least obvious of these paths:

1
Socket Access
Container has a hostPath mount of /var/run/docker.sock
→
2
Create a Container
Uses the socket to start a new, adversary-controlled container
→
3
Mount the Host
New container mounts the host filesystem
→
4
Full Host Compromise
Effective root, without ever technically "escaping" the original container

Worked Example

Worked Example:

A quarterly pod-spec sweep in a mid-size cluster turns up a monitoring agent DaemonSet running with hostPID: true and the SYS_PTRACE capability, both genuinely required for it to inspect host process metrics. Alongside it, an unrelated internal tooling deployment, added eight months earlier by a different team, is running with privileged: true and a hostPath mount of /var/run/docker.sock, with no comment or ticket reference explaining why. The monitoring agent is a known, reviewed exception; the tooling deployment is not, and represents an unreviewed full host-escape path sitting in production. The hunt's value was not detecting an active exploit: it was finding the door before anyone walked through it.

Image and Registry Hunting

Why Image Provenance Matters

A container is only as trustworthy as the image it was built from. T1610 (Deploy Container) and T1612 (Build Image on Host) both center on getting an adversary-influenced image running, whether pulled from an untrusted registry or built directly against the runtime's API to skip registry-based scanning entirely.

Four Hunts to Run

  • Hunt for images pulled from registries outside the organization's approved allow-list. A workload pulling directly from Docker Hub's public namespace, or from a registry that has never appeared in the cluster's image-pull history before, is a lead worth checking, especially for images with generic, imitation-prone names (nginx-official, redis-cache) that mimic legitimate ones.
  • Hunt for images without a corresponding CI/CD provenance record. A mature pipeline signs or attests every image it builds. A running workload backed by an image with no matching build record, commit hash, or signature is either a manual deploy that bypassed the pipeline (a process gap worth closing) or something more concerning.
  • Hunt for direct daemon/API build requests (T1612). A build request sent straight to the Docker or containerd API, rather than through the CI/CD pipeline's normal build step, bypasses every image-scanning gate that only inspects images at registry-push time.
  • Watch for resource-consumption anomalies consistent with cryptomining. A generic-looking image that spikes CPU utilization to near 100% shortly after deployment, with no corresponding legitimate workload change, is one of the most common real-world outcomes of a compromised or malicious container image: financially motivated actors overwhelmingly prioritize compute theft over data theft in container environments, because it monetizes instantly and often goes unnoticed for weeks in a busy cluster.
Note: Image scanning at admission time (via an admission controller enforcing signature verification or vulnerability thresholds) prevents known-bad images from ever running. Hunting is what covers everything admission-time scanning cannot: images that were clean at scan time but whose registry was compromised afterward, and behavior-based abuse (T1612) that never goes through the registry at all.

Building Container and Kubernetes Hunt Hypotheses

Applying ABLE to Containers

The same ABLE structure and PEAK hunt-plan template from earlier chapters applies directly, with the Kubernetes audit log and Falco runtime alerts as the primary data sources in place of Sysmon.

Hunt Plan Template

TEMPLATE
HUNT PLAN
Analyst: [name]
Hypothesis: An adversary with a stolen or over-permissioned service account
  token used the Kubernetes API to exec into a running pod and enumerate
  secrets across namespaces it should not have access to (T1609, T1613,
  T1552.007)
ATT&CK: T1609 - Container Administration Command

Scope:
- Cluster: production-eks-01, all namespaces
- Time window: Last 30 days
- Data sources: Kubernetes API server audit log, Falco alerts, RBAC bindings

Expected evidence if hypothesis TRUE:
- Audit log: pods/exec subresource calls from a service account identity
  (not a human SSO user) outside its owning namespace
- Audit log: get/list against the secrets resource across multiple
  namespaces from the same identity in a short window
- No corresponding CI/CD job or documented debugging ticket for the activity

Expected evidence if hypothesis FALSE:
- All cross-namespace secret access traces to a documented, RBAC-scoped
  operator/controller with a legitimate multi-namespace function

Planned queries:
1. Filter audit log for objectRef.subresource == "exec" grouped by user.username
2. Cross-reference identities with type service-account against their
   owning namespace vs. the namespace of the objectRef
3. Filter audit log for verb in (get, list) against resource == "secrets",
   group by user, count distinct namespaces touched
4. Cross-reference flagged identities against the cluster's RoleBinding/
   ClusterRoleBinding definitions to confirm intended scope

Three Ready-to-Adapt Hypotheses

For a first pass at a container/Kubernetes hunt program:

  • Privileged pod sweep: "A workload in this cluster is running with privileged: true, a sensitive hostPath mount, or hostPID/hostNetwork without a documented exception." (T1611). Test via a periodic pod-spec sweep, not a log query.
  • RBAC self-escalation: "An adversary with limited namespace access created or modified a ClusterRoleBinding to grant themselves or a controlled service account cluster-admin." (T1078.004-adjacent, container-specific privilege escalation). Test against the audit log filtered on RBAC resource verbs.
  • Cryptomining via unauthorized deploy: "An adversary used T1610 to deploy a container from an unapproved registry that consumes anomalous CPU/GPU resources shortly after creation." Test by correlating the audit log's pod-create events against a resource-utilization baseline per namespace.

Key Takeaways

  • Containers share a kernel with no hardware-enforced isolation boundary, and the orchestrator's API is itself a cluster-wide attack surface: a compromised credential with API access can be functionally equivalent to domain admin.
  • ATT&CK for Containers anchors container hunt hypotheses around five techniques: T1610 (Deploy Container), T1611 (Escape to Host), T1612 (Build Image on Host), T1613 (Container and Resource Discovery), and T1552.007 (Unsecured Credentials: Container API).
  • The Kubernetes API server audit log (audit.k8s.io/v1) is the primary hunting surface: pods/exec subresource calls, system:anonymous access, and RBAC self-escalation via ClusterRoleBinding creation are the three highest-value hunt patterns.
  • Falco fills the gap the audit log cannot: syscall-level runtime behavior inside a container, such as an unexpected shell spawn, a sensitive file write, or unexpected egress.
  • Most real container escapes exploit an over-permissioned pod spec (privileged: true, a sensitive hostPath mount, hostPID/hostNetwork) rather than a kernel zero-day, making a periodic pod-spec sweep one of the highest-value hunts in this chapter.
  • Financially motivated container compromise overwhelmingly trends toward cryptomining rather than data theft, because compute theft monetizes immediately and can go unnoticed in a busy cluster for weeks.

Knowledge Check

Click an answer to reveal the explanation.

Q1. Why is a container escape (T1611) fundamentally different from a VM escape?

  • A. Container escapes are always network-based, while VM escapes require physical access
  • B. Containers share the host kernel with no hardware-enforced isolation boundary, so a successful escape reaches the host directly rather than crossing a hypervisor boundary
  • C. Container escapes only affect the container itself and never reach the host
  • D. There is no meaningful difference between the two

Q2. A hunter finds a pod spec running with hostPath mounting /var/run/docker.sock. What does this grant the container?

  • A. Read-only access to container logs
  • B. Effective control over the container runtime itself, equivalent to root access on the host
  • C. Access only to that container's own metadata
  • D. Nothing beyond normal container isolation, since Docker socket access is sandboxed

Q3. What does the objectRef.subresource: "exec" field in a Kubernetes audit log entry indicate?

  • A. A new pod was scheduled onto a node
  • B. Someone opened an interactive session inside a running pod, equivalent to kubectl exec
  • C. A ConfigMap was updated
  • D. A node was cordoned for maintenance

Q4. Why does Falco's "Terminal shell in container" rule have disproportionate hunting value?

  • A. Because it detects every possible container attack technique in one rule
  • B. Because most production application containers are built without an interactive-shell need, so a shell spawning inside one is rarely legitimate and frequently indicates a working exploit
  • C. Because it only fires on privileged containers, which are always malicious
  • D. Because it replaces the need for a Kubernetes audit log entirely

Q5. What is the primary purpose of a periodic pod-spec sweep, as distinct from an audit-log hunt?

  • A. To detect malware signatures inside running container images
  • B. To find static configuration risks (privileged mode, sensitive hostPath mounts, hostPID/hostNetwork) that create an escape path before any exploit ever runs, since these are facts about the spec rather than events in a log
  • C. To replace the need for RBAC entirely
  • D. To measure cluster CPU and memory utilization for capacity planning