CHAPTER 03 35 MIN READ INTERMEDIATE

Dynamic Analysis & Sandboxing

Chapter 2 covered pulling a sample apart without ever running it: hashes, PE structure, strings, imports, packer signatures. That approach has a hard ceiling, because a sufficiently packed or obfuscated binary can hide its real logic from every one of those static techniques. Dynamic analysis clears that ceiling by a different route entirely: execute the sample in a controlled environment and watch what it actually does. The trade is real. You see genuine behavior instead of a guess about behavior, but only the specific code path the sample happens to take during that run, nothing more.

sandboxing API monitoring network capture behavioral indicators

Why Dynamic Analysis Exists Alongside Static Analysis

Static analysis answers "what could this file do." Dynamic analysis answers "what did this file actually do, this one time, on this one run." Those are different questions, and a mature analysis workflow needs both because each one covers the other's blind spot.

The Blind Spot Each One Covers

A packer that encrypts the real payload until runtime defeats string extraction and import table analysis almost completely, since the interesting code simply isn't sitting in the file in a readable form when a disassembler opens it. Run that same packed sample, though, and the unpacking stub does the work for you: it decrypts the payload into memory and hands control to it, which is exactly the moment dynamic analysis starts paying off.

The Cost: Only One Code Path Per Run

A static analyst looking at a decompiled function sees every branch that function could take. A dynamic analyst watching a live execution sees only the branch the program actually took on that run, with that input, in that environment, at that moment. If a piece of malware checks the date and only detonates its ransomware payload after a certain point, or waits for a command from its command-and-control server before doing anything destructive, a five-minute sandbox run that never receives that command will look almost boring. The sample executed. It didn't do much. That's not proof the sample is harmless, it's proof that one particular run didn't trigger one particular code path.

Complementary, Not Competing

Static analysis maps out what's possible; dynamic analysis confirms what's real, at least for the paths it happened to exercise. An analyst who treats a single sandbox run as a complete verdict on a sample is making the same mistake as an analyst who treats a strings dump as a complete verdict, just from the opposite direction.

Note: Multiple dynamic analysis runs, with different simulated dates, different simulated user interaction, or different network conditions, surface more of a sample's conditional behavior than a single run ever will. Treat one run as a sample, not a census.

Sandbox Environment Design

Isolated, Revertible VM

The foundation of safe dynamic analysis is an isolated virtual machine that can be reverted to a known-clean state in seconds. Without that revert step, every subsequent analysis run happens on a progressively more compromised, less trustworthy baseline, and eventually the analyst can no longer tell which artifacts came from the current sample and which were left behind by an earlier one.

1
Build the VM
Set up the isolated analysis machine.
→
2
Install Monitoring Tools
Add whatever the analysis needs: ProcMon, API monitors, Wireshark.
→
3
Snapshot
Capture that exact clean state before anything runs.
→
4
Execute the Sample
Run the sample and observe.
→
5
Revert
Discard whatever the malware changed, dropped, or corrupted, back to the clean snapshot.

Simulated Internet

Isolation extends past "it's a VM" to "it has no real path to the internet." Malware that reaches out to its actual command-and-control infrastructure from an analysis environment can tip off the operator that their sample is under scrutiny, and in the worst case it can use that real connectivity to cause harm well outside the lab. The fix most labs use is a controlled, simulated internet: tools like INetSim or FakeNet-NG stand in for DNS, HTTP, HTTPS, and other common services, answering the malware's requests convincingly enough that the sample believes it has genuine network access. That matters because a lot of malware's interesting behavior, its actual C2 logic, its beaconing pattern, the second-stage payload it tries to pull down, only triggers once the sample believes a network response came back. A sample that gets nothing but connection failures for every request it makes may simply give up and go quiet, and the analyst walks away thinking the sample doesn't do much of anything.

Automated vs Manual Analysis

Within that isolated design, there are two broad ways to actually run the analysis.

Fully Automated SandboxManual, Analyst-Driven Analysis
ExamplesCuckoo Sandbox (self-hosted), ANY.RUN, Joe Sandbox, Hybrid AnalysisAnalyst working directly inside the isolated VM
What it doesExecutes the sample and captures process activity, API calls, file and registry changes, and network traffic automatically, then hands back a structured reportAnalyst decides when to interact with the sample, feed it fake network responses, or pause and inspect a running process
StrengthFast and consistent; the right default for triage volumeCan respond to what the malware does in real time rather than waiting for a fixed observation window
Best forHigh-volume triageSamples that behave differently under interaction, need a specific file or registry key present first, or were already flagged by an automated report as worth a closer look
Warning: Never run dynamic analysis on a machine, physical or virtual, that has a real network path to your production environment or the open internet. An isolated VM with no outbound connectivity, paired with a simulated internet like INetSim or FakeNet-NG, is the baseline, not an optional precaution. A sample that reaches real C2 infrastructure from your analysis network can alert the attacker and, depending on its capabilities, cause damage that has nothing to do with the analysis itself.

Process and API Call Monitoring

Process-Level Visibility: ProcMon and Process Hacker

Once the sample is executing, the first layer of visibility is process-level activity: what got created, what files and registry keys got touched, and what handles and threads a process is holding. Process Monitor, universally called ProcMon, is the standard tool for this in a Windows analysis environment. It logs file system activity, registry activity, and process and thread activity in real time, with filtering that lets an analyst narrow a firehose of system noise down to just the events tied to the sample's process tree. Process Hacker and Process Explorer serve an adjacent purpose: real-time visibility into what's actually running, parent-child process relationships, loaded modules, open handles, and network connections per process, which is often the fastest way to spot a sample spawning an unexpected child process or injecting into one that's already running.

API-Level Visibility: Beyond the Import Table

ProcMon and Process Hacker both work at the level of system calls and process metadata. API monitoring goes a layer deeper, showing the actual sequence of Windows API calls a sample makes as it runs, tools like API Monitor, or hooking-based instrumentation built for this purpose, intercept calls into system DLLs and log them with their arguments and return values. This is a meaningfully more granular signal than the static import table from Chapter 2. The import table tells you a sample links against VirtualAlloc, WriteProcessMemory, and CreateRemoteThread. It doesn't tell you the order those calls happen in, what arguments get passed, or whether they even execute at all during a given run. API monitoring shows the actual call sequence, which is often enough on its own to recognize a process injection pattern, a classic allocate-write-execute sequence looks nothing like a case where the same imports exist in the binary but are never actually invoked.

Tip: Reading ProcMon output takes a little practice, mostly around learning to filter out the baseline noise every Windows process generates. A malicious sample's activity sits inside thousands of lines of ordinary registry reads and DLL loads that have nothing to do with its behavior, and the skill is narrowing the view to the process of interest and the operations that actually matter (file writes, registry value creation, new process launches).
PROCMON (illustrative)
12:04:17.9021  sample.exe   3812  RegSetValue    HKCU\Software\Microsoft\Windows\CurrentVersion\Run\Updater   SUCCESS
12:04:18.0114  sample.exe   3812  CreateFile     C:\Users\analyst\AppData\Roaming\svchost32.exe               SUCCESS
12:04:18.0347  sample.exe   3812  WriteFile      C:\Users\analyst\AppData\Roaming\svchost32.exe               SUCCESS
12:04:18.2209  sample.exe   3812  Process Create C:\Users\analyst\AppData\Roaming\svchost32.exe                SUCCESS
12:04:18.4432  svchost32.exe 4108 VirtualAllocEx explorer.exe (PID 1988)                                        SUCCESS
12:04:18.4501  svchost32.exe 4108 WriteProcessMemory explorer.exe (PID 1988)                                    SUCCESS
12:04:18.4587  svchost32.exe 4108 CreateRemoteThread explorer.exe (PID 1988)                                    SUCCESS

That short sequence alone tells a real story: a dropped file disguised with a system-sounding name, a Run key set for persistence, and then a textbook process injection into a legitimate process. None of that specific sequence, in that specific order, is visible from a static import table.

Network Traffic Capture

Capturing Traffic: Wireshark and PCAP

Alongside process behavior, the sample's network activity is one of the most immediately actionable outputs of a dynamic run. Wireshark, run against the sandbox's virtual network interface, or a sandbox platform's built-in PCAP capture, records everything the sample sends and receives: DNS requests resolving a C2 domain, HTTP or HTTPS beaconing to a callback server, or raw traffic following the sample's own custom protocol rather than a standard web protocol at all. Even a simulated internet environment like INetSim or FakeNet-NG that answers the sample's requests still lets an analyst capture the requests themselves, which domains it tried to resolve, what paths it requested over HTTP, what data it tried to send outbound.

The Encryption Problem

A growing share of C2 traffic runs over TLS, the same as ordinary web traffic, and a raw packet capture of encrypted traffic shows connection metadata without the content that would tell an analyst much about what's actually being communicated. Analysts get around this two ways.

ApproachHow It WorksLimitation
TLS interceptionInstall a trusted proxy certificate in the analysis VM and route the sample's traffic through a controlled proxy, so outbound TLS connections terminate at the proxy and the plaintext becomes visible before it's re-encrypted onward or droppedWorks cleanly against malware that doesn't validate the certificate it's handed, which describes a large share of commodity samples, but fails against anything doing certificate pinning
SNI, timing, and size fallbackRead what's still visible without decrypting anything: the SNI field in the TLS handshake exposes the destination domain even when the payload is encryptedNo visibility into actual content, but connection timing and payload size patterns can still reveal a beaconing interval
Tip: Even when TLS interception isn't possible or the malware detects and refuses the proxy, capture the traffic anyway. SNI values, DNS queries, destination IPs, and beacon timing are all recoverable from an encrypted capture and are often enough on their own to extract usable C2 indicators.
PCAP SUMMARY (illustrative)
No.   Time      Source        Destination     Protocol  Info
1     0.0012    10.10.10.15   10.10.10.2       DNS       Standard query A cdn-update-service[.]net
2     0.0341    10.10.10.15   198.51.100.44    TCP       443 -> 51221 [SYN, ACK]
3     0.0389    10.10.10.15   198.51.100.44    TLSv1.2   Client Hello, SNI=cdn-update-service[.]net
4     0.0612    198.51.100.44 10.10.10.15      TLSv1.2   Server Hello, Certificate
5     60.0021   10.10.10.15   198.51.100.44    TLSv1.2   Application Data (encrypted, 312 bytes)
6     120.0034  10.10.10.15   198.51.100.44    TLSv1.2   Application Data (encrypted, 308 bytes)

The 60-second gap between application data frames in that summary is the tell, even without decrypting a single byte of payload: a beacon interval this regular almost never comes from ordinary user browsing.

Registry and Filesystem Monitoring as a Persistence Signal

Everything a sample creates, modifies, or deletes on disk and in the registry during a dynamic run is a concrete, checkable artifact, and it's worth treating registry and filesystem monitoring as its own dedicated lens rather than a byproduct of watching ProcMon scroll by.

What to Watch

  • Autorun keys under Run or RunOnce.
  • Scheduled tasks created to relaunch the sample after reboot.
  • Services registered to start automatically.
  • Dropped files in user or system directories, often under an innocuous-sounding name meant to blend into a directory listing.

Why These Artifacts Matter

These artifacts do double duty. In the moment, they explain how the sample intends to survive a reboot or keep running, which is directly useful for scoping and remediation during an incident. Longer term, they become indicators of compromise: a specific registry value name, a specific dropped file path and hash, a specific scheduled task name. Chapter 6 covers turning artifacts exactly like these into YARA and Sigma rules, so the persistence mechanism a sample uses during dynamic analysis isn't just descriptive information, it's the raw material for a detection that catches the next sample using the same technique.

Catching Secondary Payloads

Filesystem monitoring also catches something static analysis structurally cannot: secondary payloads. A downloader or loader sample's entire job is often to retrieve and drop a second file, and that second file only exists on disk once the first sample actually runs and fetches it. Static analysis of the original sample sees a network call and maybe a URL in a string. Dynamic analysis sees the actual file that call produces, sitting on disk, ready to be hashed and analyzed in its own right.

Behavioral Indicators as Ground Truth

Pull together everything a dynamic run surfaces and what you end up with is a set of behavioral indicators that static analysis alone cannot produce, because each one only exists once the sample actually executes.

What counts as ground-truth behavioral evidence:
  • C2 domains and IP addresses the sample actually contacted, as opposed to ones merely present as strings in the binary.
  • Dropped files and their hashes, rather than a guess based on an embedded resource.
  • The exact registry keys it set for persistence.
  • The specific injection target process, and the API sequence used to get code into it.
  • Mutex and named-pipe names a sample creates purely to mark a machine as already infected, so it doesn't reinfect the same host twice. A unique mutex or pipe name is often one of the most reliable single indicators a family produces.
Tool / PlatformWhat It MonitorsTypical Use
Process Monitor (ProcMon)File system, registry, process, and thread activity in real timeTracing exactly what a running sample touches on disk and in the registry
Process Hacker / Process ExplorerLive process tree, loaded modules, open handles, per-process network connectionsSpotting unexpected child processes, injected modules, and live connections
API Monitor / hooking instrumentationSequence of Windows API calls with arguments and return valuesConfirming call order for things like process injection, beyond what the import table shows
WiresharkRaw network traffic: DNS, HTTP/HTTPS, custom protocolsCapturing C2 domains, beacon timing, and payload size patterns
Cuckoo Sandbox / ANY.RUN / Joe Sandbox / Hybrid AnalysisAutomated end-to-end capture: process, API, file, registry, and network activityFast, consistent triage across a high volume of samples

This set of indicators is what feeds directly into the detection engineering work in Chapter 6. A YARA rule built purely from static strings can be fragile against the next minor variant of a family. A Sigma rule built from the actual process injection sequence, the actual persistence key, or the actual mutex name observed during dynamic analysis targets the behavior a family relies on, which tends to survive across variants far better than any single string does.

A Caution: Sandbox-Aware Malware

Everything in this chapter assumes the sample behaves the same way in the sandbox as it would on a real victim machine, and that assumption doesn't always hold.

What Sandbox-Aware Malware Checks For

  • Specific artifacts left by common hypervisors.
  • The presence of analysis tools running on the system.
  • How the machine responds to certain timing checks.

Detecting any of these, the sample either refuses to execute its actual malicious payload or deliberately behaves differently. A sample that goes quiet in the sandbox isn't necessarily benign; it may just be smart enough to know it's being watched.

This chapter deliberately doesn't go deep on how that detection works or how analysts work around it. That's a big enough topic, and a specialized enough one, to earn its own dedicated treatment in Chapter 8, Advanced Evasion and Anti-Analysis. For now, the important takeaway is simpler: treat an unremarkable sandbox report as one data point, not a clean bill of health, especially for samples that look otherwise sophisticated or that a static pass already flagged as suspicious.

Note: A boring dynamic analysis report on a sample that static analysis flagged as heavily obfuscated or packed is itself a signal worth paying attention to. It's exactly the pattern a sandbox-aware sample produces, and Chapter 8 covers the specific techniques behind it.

Key Takeaways

  • Dynamic analysis executes a sample to observe real behavior, revealing what packing and obfuscation hide from static analysis, but it only shows the specific code path the sample takes during that run, not every path it's capable of.
  • A safe sandbox is an isolated, revertible VM combined with a simulated internet like INetSim or FakeNet-NG, so malware believes it has network access and its C2 logic actually triggers, without any real path out to the internet.
  • Fully automated sandboxes (Cuckoo Sandbox, ANY.RUN, Joe Sandbox, Hybrid Analysis) are fast and consistent for triage volume; manual, analyst-driven dynamic analysis trades speed for the ability to interact with and respond to a sample in real time.
  • ProcMon and Process Hacker show process, file, and registry activity; API monitoring goes deeper, showing the actual sequence and arguments of Windows API calls, which is more granular than a static import table alone.
  • Network capture surfaces C2 domains and beaconing patterns even against encrypted traffic, through TLS interception where the sample doesn't pin certificates, or through SNI, timing, and size analysis when decryption isn't possible.
  • Behavioral indicators, actual C2 infrastructure, dropped file hashes, persistence keys, injection targets, and mutex or pipe names, are ground truth that static analysis alone can't produce, and they feed directly into the YARA and Sigma rule writing in Chapter 6. Some samples detect sandboxing and suppress their real behavior, a limitation Chapter 8 covers in depth.

Knowledge Check

Click an answer to reveal the explanation.

An analyst runs a sample in a sandbox with no network access at all and observes almost no interesting behavior. What's the most likely explanation, and what should the analyst try next?

A lot of malware's interesting behavior only executes once it believes it has received a network response, whether that's a C2 command, a second-stage payload, or simple DNS resolution succeeding. A sandbox with zero network access, real or simulated, often produces a quiet, unremarkable run for exactly that reason. INetSim or FakeNet-NG let the sample believe it has genuine connectivity so that logic actually fires.

During dynamic analysis, ProcMon shows a sample calling VirtualAllocEx, WriteProcessMemory, and CreateRemoteThread in sequence, all targeting explorer.exe. Why is this more useful than knowing the sample's static import table includes those same three functions?

A static import table tells you a binary links against a function, nothing about whether that function is ever actually called, in what order, or with what arguments. Seeing the allocate-write-execute sequence actually occur, and against a specific target process, confirms real process injection behavior rather than a theoretical capability.

A sandbox captures a sample's network traffic, but it's all TLS-encrypted and the sample pins its certificate, so a proxy-based TLS interception attempt fails. What can the analyst still extract from the capture?

Certificate pinning defeats proxy-based TLS interception because the sample refuses to trust anything but the certificate it expects. That doesn't blind the analyst completely, though: the SNI field in the Client Hello reveals the destination domain in plaintext, and connection timing and payload size are visible at the network layer regardless of encryption, which is often enough to identify a beaconing pattern and extract a usable indicator.