Container and Kubernetes Hunting
A container is not a lightweight virtual machine. It is a set of Linux namespaces and cgroups sharing one kernel with every other container on the host, orchestrated by a control plane that itself is reachable over the network as an API. That combination creates attack surface no endpoint-hunting chapter in this course has touched: escaping a container means reaching the host kernel directly, and compromising one over-privileged pod's service account can mean compromising the entire cluster through the Kubernetes API. This chapter builds the hunting model for that environment.
Why Containers Need a Different Hunting Model
Every technique covered in Chapters 1 through 9 assumes a single operating system instance with its own kernel, its own process tree, and Sysmon or an EDR agent watching it directly. Kubernetes breaks that model in three specific ways that change what a hunter should look for and where.
Three Ways Kubernetes Breaks the Model
- Containers share a kernel. Unlike a hypervisor-based VM, a container has no hardware-enforced isolation boundary. A kernel vulnerability, a misconfigured capability, or a mounted host resource can let a process inside a container directly affect the host or every other container running on it. This is the entire premise of T1611, Escape to Host.
- The orchestrator is itself an API-reachable attack surface. The Kubernetes API server is a control plane that can create, delete, and exec into workloads across the entire cluster. A credential that grants API access (a stolen kubeconfig, an over-permissioned service account token) is functionally equivalent to domain admin in a Windows environment, except the "domain" is every workload the cluster runs.
- Workloads are ephemeral by design. A compromised pod might live for minutes before the orchestrator reschedules it. Traditional incident response assumptions (the compromised host is still there tomorrow for forensic imaging) do not hold. Hunting has to work from the audit trail and runtime telemetry, because the pod itself may already be gone.
Applying the Same ATT&CK Discipline
This chapter maps to the same ATT&CK-driven hypothesis discipline from Chapter 2, using the ATT&CK for Containers matrix published by the Center for Threat-Informed Defense, plus the container-relevant sub-techniques already present in TTPHUNT: T1053.007 (Container Orchestration Job) and T1059.013 (Container CLI/API). Where this chapter goes further is the cluster-level control plane: RBAC, the API server audit log, and container runtime behavior.
ATT&CK for Containers Overview
The Center for Threat-Informed Defense published ATT&CK for Containers as an extension to the Enterprise matrix, adding techniques specific to container and orchestration abuse rather than forcing container attacks into host-centric technique definitions that did not quite fit.
Five Anchor Techniques
| Technique | What It Covers | Primary Hunt Surface |
|---|---|---|
| T1610: Deploy Container | Adversary deploys a new container (often from a malicious or unmodified public image) to execute code, establish persistence, or evade defenses that were watching the existing workloads | Kubernetes audit log: create verb on pods/deployments from an unexpected principal or CI/CD identity outside its normal deploy window |
| T1611: Escape to Host | Adversary breaks the container boundary to execute code on the underlying node, gaining access to every other container and the host's own credentials | Privileged pod specs, mounted hostPath/Docker socket, runtime (Falco) syscall detections for namespace/capability abuse |
| T1612: Build Image on Host | Adversary sends a build request directly to the container runtime's API (e.g., the Docker daemon socket) to build and run an image locally, bypassing defenses that only scan images pulled from a registry | Docker/containerd API build-request logs; runtime detection for unexpected daemon socket access |
| T1613: Container and Resource Discovery | Adversary enumerates containers, pods, namespaces, and cluster resources to understand what they have landed in and what else is reachable | Audit log list/get verbs against pods, namespaces, secrets at unusual volume from a single identity |
| T1552.007: Unsecured Credentials: Container API | Adversary queries the Docker API or Kubernetes API (frequently via an over-permissioned pod service account token) to retrieve credentials, secrets, or logs containing sensitive data | Audit log access to the secrets resource; service account token usage from a pod that has no legitimate reason to read secrets outside its own namespace |
kubectl exec into a running pod, or uses the Docker CLI/API directly, to execute commands. It requires no exploit and no misconfiguration beyond overly broad access, which is exactly why pods/exec auditing (Section 3) is one of the highest-value hunts in this chapter.
Kubernetes Audit Log Hunting
The Audit Log as Primary Data Source
The Kubernetes API server can emit a structured audit log (audit.k8s.io/v1) for every request it processes, if an audit policy is configured. This is the single most important data source in this chapter: every kubectl command, every controller reconciliation, and every service account API call passes through the API server and can be captured here.
Core fields to build hunt queries around: verb (get, list, create, update, patch, delete, deletecollection), objectRef (the resource type, namespace, and name being acted on, including a subresource field for things like exec and portforward), user (the identity making the call, a human via SSO, or a service account), and sourceIPs.
Example Audit Log Entry
{
"apiVersion": "audit.k8s.io/v1",
"kind": "Event",
"level": "Metadata",
"verb": "create",
"objectRef": {
"resource": "pods",
"namespace": "production",
"name": "web-frontend-7c9f8d-xk2lp",
"apiVersion": "v1",
"subresource": "exec"
},
"user": { "username": "jane.doe@corp.com", "groups": ["system:authenticated"] },
"sourceIPs": ["203.0.113.45"],
"requestURI": "/api/v1/namespaces/production/pods/web-frontend-7c9f8d-xk2lp/exec?command=/bin/sh"
}
Three Hunt Patterns
These cover most of the value in this log source:
- Exec into production pods. Filter for
objectRef.subresource: "exec"or"attach"against namespaces that should never need interactive debugging (production, especially anything handling regulated data). A legitimate on-call engineer execs in occasionally; a pattern of repeated execs across many pods from one identity in a short window looks like discovery or lateral movement, not debugging. - Anonymous or unauthenticated API access. A
user.usernameofsystem:anonymoussucceeding against any resource indicates the API server (or an exposed component like the kubelet's read-only port) is reachable without authentication, a critical exposure regardless of what specific resource was accessed. - RBAC self-escalation. Watch for
create/updateverbs againstclusterrolebindings,rolebindings, orclusterrolesperformed by an identity that is not a documented cluster administrator. Binding oneself (or a controlled service account) tocluster-adminis the single most direct privilege escalation path in Kubernetes.
Metadata level, may exclude the request/response bodies you need for some investigations. Before trusting a hunt result of "no matches," verify the audit policy actually covers the verbs, resources, and namespaces in scope: the same Equip-stage discipline from Chapter 3's SANS PAM coverage.
Runtime Detection with Falco
Why Falco Fills the Gap
The Kubernetes audit log tells you what was requested through the API. It has no visibility into what a process does once it is already running inside a container: a reverse shell spawned by an exploited application, a write to a sensitive host-mounted path, or a process reading the pod's own service account token to exfiltrate it. That gap is covered by syscall-level runtime security tooling, of which Falco (a CNCF graduated project) is the de facto standard.
Falco rules match on syscall events using a condition expression built from fields like proc.name, container.id, fd.name, and evt.type, reusable via macros and lists. The canonical example, "Terminal shell in container," illustrates the pattern every container runtime hunt should understand:
Example Falco Rule
- rule: Terminal shell in container
desc: Detect a shell being spawned in a container
condition: > spawned_process and container and shell_procs and proc.tty != 0
output: > Shell spawned in container (user=%user.name
container_id=%container.id container_name=%container.name
shell=%proc.name parent=%proc.pname cmdline=%proc.cmdline)
priority: WARNING
This single rule is disproportionately valuable because most production container images are built without an interactive shell need in mind: a shell spawning inside a minimal application container, attached to a TTY, is rarely a legitimate operational action and very frequently the first sign of a working remote code execution exploit.
More Falco Hunt Patterns
| Falco Hunt Pattern | What It Catches |
|---|---|
| Shell spawned in a container with no interactive-shell business need | Post-exploitation shell from a web application RCE, T1609 Container Administration Command abuse |
Write below /etc, sensitive mount paths, or a read-only root filesystem attempt |
Persistence via config tampering; also flags containers that were supposed to run with a read-only root filesystem but do not |
| Outbound connection to an unexpected port/destination from a container that should have no egress need | C2 callback, data exfiltration from a workload with a narrow expected network profile |
| Process attempting to read the container's own service account token file outside expected client libraries | Credential harvesting consistent with T1552.007, Unsecured Credentials: Container API |
Container Escape and Privileged Pod Hunting
Why Pod Specs Matter More Than Exploits
Most real-world container escapes (T1611) do not require a kernel zero-day. They require a pod spec that was granted more access than the workload needed, deployed by a team optimizing for "it works" over least privilege. Hunting for escape risk is largely a matter of hunting the pod specification itself, before any exploit ever runs.
Escape-Enabling Pod Spec Settings
| Pod Spec Setting | Why It Enables Escape |
|---|---|
securityContext.privileged: true |
Disables nearly all container isolation, granting access to all host devices and effectively equivalent to root on the host |
hostPath volume mount, especially of /, /etc, or /var/run/docker.sock |
Gives the container direct read/write access to the host filesystem or, in the case of the Docker socket, full control over the container runtime itself (equivalent to root on the host) |
hostNetwork: true |
Removes network namespace isolation, exposing the container to every network interface and service the host itself can reach, and vice versa |
hostPID: true |
Exposes the host's entire process namespace to the container, allowing it to see and potentially signal/ptrace processes running outside its own namespace |
Added Linux capabilities (SYS_ADMIN, SYS_PTRACE, SYS_MODULE) |
Each grants a specific piece of root-equivalent kernel functionality; SYS_ADMIN in particular exposes a large surface of historically exploitable escape primitives |
Hunting the Spec, Not the Log
Because these are static configuration facts rather than runtime events, the hunt is a periodic sweep, not a log query: enumerate every pod spec in the cluster (via kubectl get pods -o json across all namespaces, or an admission-controller policy engine's audit mode) and flag any workload carrying one of the settings above without a documented, reviewed exception.
The Docker socket mount is worth tracing end to end, since it is one of the least obvious of these paths:
Worked Example
A quarterly pod-spec sweep in a mid-size cluster turns up a monitoring agent DaemonSet running with hostPID: true and the SYS_PTRACE capability, both genuinely required for it to inspect host process metrics. Alongside it, an unrelated internal tooling deployment, added eight months earlier by a different team, is running with privileged: true and a hostPath mount of /var/run/docker.sock, with no comment or ticket reference explaining why. The monitoring agent is a known, reviewed exception; the tooling deployment is not, and represents an unreviewed full host-escape path sitting in production. The hunt's value was not detecting an active exploit: it was finding the door before anyone walked through it.
Image and Registry Hunting
Why Image Provenance Matters
A container is only as trustworthy as the image it was built from. T1610 (Deploy Container) and T1612 (Build Image on Host) both center on getting an adversary-influenced image running, whether pulled from an untrusted registry or built directly against the runtime's API to skip registry-based scanning entirely.
Four Hunts to Run
- Hunt for images pulled from registries outside the organization's approved allow-list. A workload pulling directly from Docker Hub's public namespace, or from a registry that has never appeared in the cluster's image-pull history before, is a lead worth checking, especially for images with generic, imitation-prone names (
nginx-official,redis-cache) that mimic legitimate ones. - Hunt for images without a corresponding CI/CD provenance record. A mature pipeline signs or attests every image it builds. A running workload backed by an image with no matching build record, commit hash, or signature is either a manual deploy that bypassed the pipeline (a process gap worth closing) or something more concerning.
- Hunt for direct daemon/API build requests (T1612). A build request sent straight to the Docker or containerd API, rather than through the CI/CD pipeline's normal build step, bypasses every image-scanning gate that only inspects images at registry-push time.
- Watch for resource-consumption anomalies consistent with cryptomining. A generic-looking image that spikes CPU utilization to near 100% shortly after deployment, with no corresponding legitimate workload change, is one of the most common real-world outcomes of a compromised or malicious container image: financially motivated actors overwhelmingly prioritize compute theft over data theft in container environments, because it monetizes instantly and often goes unnoticed for weeks in a busy cluster.
Building Container and Kubernetes Hunt Hypotheses
Applying ABLE to Containers
The same ABLE structure and PEAK hunt-plan template from earlier chapters applies directly, with the Kubernetes audit log and Falco runtime alerts as the primary data sources in place of Sysmon.
Hunt Plan Template
HUNT PLAN
Analyst: [name]
Hypothesis: An adversary with a stolen or over-permissioned service account
token used the Kubernetes API to exec into a running pod and enumerate
secrets across namespaces it should not have access to (T1609, T1613,
T1552.007)
ATT&CK: T1609 - Container Administration Command
Scope:
- Cluster: production-eks-01, all namespaces
- Time window: Last 30 days
- Data sources: Kubernetes API server audit log, Falco alerts, RBAC bindings
Expected evidence if hypothesis TRUE:
- Audit log: pods/exec subresource calls from a service account identity
(not a human SSO user) outside its owning namespace
- Audit log: get/list against the secrets resource across multiple
namespaces from the same identity in a short window
- No corresponding CI/CD job or documented debugging ticket for the activity
Expected evidence if hypothesis FALSE:
- All cross-namespace secret access traces to a documented, RBAC-scoped
operator/controller with a legitimate multi-namespace function
Planned queries:
1. Filter audit log for objectRef.subresource == "exec" grouped by user.username
2. Cross-reference identities with type service-account against their
owning namespace vs. the namespace of the objectRef
3. Filter audit log for verb in (get, list) against resource == "secrets",
group by user, count distinct namespaces touched
4. Cross-reference flagged identities against the cluster's RoleBinding/
ClusterRoleBinding definitions to confirm intended scope
Three Ready-to-Adapt Hypotheses
For a first pass at a container/Kubernetes hunt program:
- Privileged pod sweep: "A workload in this cluster is running with
privileged: true, a sensitivehostPathmount, orhostPID/hostNetworkwithout a documented exception." (T1611). Test via a periodic pod-spec sweep, not a log query. - RBAC self-escalation: "An adversary with limited namespace access created or modified a ClusterRoleBinding to grant themselves or a controlled service account cluster-admin." (T1078.004-adjacent, container-specific privilege escalation). Test against the audit log filtered on RBAC resource verbs.
- Cryptomining via unauthorized deploy: "An adversary used T1610 to deploy a container from an unapproved registry that consumes anomalous CPU/GPU resources shortly after creation." Test by correlating the audit log's pod-create events against a resource-utilization baseline per namespace.
Key Takeaways
- Containers share a kernel with no hardware-enforced isolation boundary, and the orchestrator's API is itself a cluster-wide attack surface: a compromised credential with API access can be functionally equivalent to domain admin.
- ATT&CK for Containers anchors container hunt hypotheses around five techniques: T1610 (Deploy Container), T1611 (Escape to Host), T1612 (Build Image on Host), T1613 (Container and Resource Discovery), and T1552.007 (Unsecured Credentials: Container API).
- The Kubernetes API server audit log (
audit.k8s.io/v1) is the primary hunting surface:pods/execsubresource calls,system:anonymousaccess, and RBAC self-escalation via ClusterRoleBinding creation are the three highest-value hunt patterns. - Falco fills the gap the audit log cannot: syscall-level runtime behavior inside a container, such as an unexpected shell spawn, a sensitive file write, or unexpected egress.
- Most real container escapes exploit an over-permissioned pod spec (
privileged: true, a sensitivehostPathmount,hostPID/hostNetwork) rather than a kernel zero-day, making a periodic pod-spec sweep one of the highest-value hunts in this chapter. - Financially motivated container compromise overwhelmingly trends toward cryptomining rather than data theft, because compute theft monetizes immediately and can go unnoticed in a busy cluster for weeks.
Knowledge Check
Click an answer to reveal the explanation.
Q1. Why is a container escape (T1611) fundamentally different from a VM escape?
Q2. A hunter finds a pod spec running with hostPath mounting /var/run/docker.sock. What does this grant the container?
Q3. What does the objectRef.subresource: "exec" field in a Kubernetes audit log entry indicate?
Q4. Why does Falco's "Terminal shell in container" rule have disproportionate hunting value?
Q5. What is the primary purpose of a periodic pod-spec sweep, as distinct from an audit-log hunt?