CHAPTER 01 25 MIN READ BEGINNER

Malware Analysis Foundations

Before you run a single tool against a real sample, you need a place to run it safely and a vocabulary for describing what it does once it runs. Most of the damage done during early malware analysis has nothing to do with missing a technical detail. It comes from analyzing a live sample on a machine that's still connected to something that matters. This chapter builds the foundation the rest of the module depends on: what malware analysis is actually for, how to classify what you're looking at, how to build an isolated lab you can detonate samples in without consequence, the difference between static and dynamic analysis, and the handling discipline that keeps the process safe and defensible.

malware classification analysis lab static vs dynamic safety and handling
New to malware analysis? This chapter assumes no reverse engineering background. It does assume you know what a process, a file system, and the Windows registry are, since almost everything malware does happens through one of those three. If any of that is unfamiliar, the Fundamentals module covers it first.

What Is Malware Analysis, and Why Does It Matter?

Malware analysis is the process of examining a malicious file or a running malicious process to determine what it is, what it does, and what it's capable of doing to the environment it lands in. That sounds simple, but the phrase hides three separate jobs that analysts actually do, often on the same sample within the same hour.

JobQuestion It AnswersPace and MethodOutput
TriageHow bad is this, and how fast do I need to act?Minutes, using surface-level static checks (Chapter 2)A verdict: known commodity loader, targeted implant, or false positive
Deep-dive analysisWhat is this sample's full capability?Slower, deliberate, combines static and dynamic techniquesA complete behavioral picture: persistence, communication, data touched, evasion
Producing usable outputWhat can the rest of the security program act on?Extracting indicators as the sample is worked, not as an afterthoughtIOCs (hash, C2 domain, mutex name, persistence registry key) and behavioral detail for detection rules

Triage: How Bad, How Fast

When a suspicious file lands on an analyst's desk, whether from an EDR alert, a user report, or an email attachment quarantined by a gateway, the first question isn't "how does this work." It's "how bad is this, and how fast do I need to act." A skilled analyst can often answer that in minutes using the surface-level static checks covered in Chapter 2.

Deep-Dive Analysis: Full Capability

Once a sample is confirmed malicious and worth the time, the goal shifts to understanding its full capability: what it persists as, what it communicates with, what data it touches, and how it evades detection. Static and dynamic techniques get combined here to build a complete picture of behavior rather than a quick verdict.

Producing Usable Output: The Real Product

Analysis that stays in an analyst's head helps nobody. The real product of malware analysis is indicators of compromise (IOCs) that other analysts can search for, and behavioral detail that a detection engineer can turn into a rule. A hash, a C2 domain, a mutex name, a registry key used for persistence: each of these, extracted correctly, becomes something a SOC can hunt for across an entire environment. Chapter 6 covers turning that output into YARA and Sigma rules directly, but the discipline of extracting clean, accurate indicators starts here, in how you approach the sample in the first place.

Note: Every chapter after this one assumes the goal is never "understand the malware for its own sake." The goal is always to produce something actionable: a verdict, an indicator, or a detection rule. Analysis that doesn't feed one of those three outcomes is an academic exercise, not a security function.

Malware Classification by Type

Classifying malware by type gives analysts a shared vocabulary and a starting set of expectations about behavior. Each category below describes a distinct strategy for causing harm, spreading, or maintaining access, though as the closing note in this section explains, real samples rarely respect these boundaries cleanly.

Glossary, malware types:
  • Virus: attaches itself to a legitimate host file or program and relies on that host being executed or shared to spread. It typically needs some form of user action, running the infected program, opening the infected document, to propagate.
  • Worm: self-propagating. It spreads across a network on its own, exploiting a vulnerability or a weak configuration to move from host to host without requiring a user to run anything. Host-dependent versus self-propagating is the classic line drawn between virus and worm, even though both terms get used loosely in casual conversation.
  • Trojan: disguises itself as something legitimate or desirable to trick a user into running it voluntarily. Unlike a virus, it doesn't attach to a host file, and unlike a worm, it doesn't spread on its own. It depends entirely on social engineering to get executed the first time.
  • Ransomware: encrypts a victim's files and demands payment for the decryption key. Modern ransomware operations frequently add data theft and extortion on top of encryption.
  • Backdoor: gives an attacker persistent remote access to a compromised system, often installed as a secondary payload after initial access.
  • Rootkit: focuses specifically on hiding, modifying the operating system's own reporting mechanisms so that its files, processes, or network connections don't show up in normal system views.
  • Loader / dropper: whose entire job is to fetch or unpack a second-stage payload and execute it, functioning as the delivery mechanism rather than the final capability.
  • Spyware / infostealer: focuses on collecting information, credentials, browser data, cryptocurrency wallet files, and exfiltrating it, usually without any destructive action that would tip off the victim.
TypePrimary GoalTypical Persistence
VirusSpread via host file infectionEmbedded in infected files
WormSelf-propagate across a networkOften none beyond active spread
TrojanGain execution via deceptionVaries by payload delivered
RansomwareExtort via encryption and data theftShort-lived, runs to completion
BackdoorMaintain remote accessRegistry run keys, scheduled tasks, services
RootkitHide presence from the OSKernel or boot-level hooks
Loader / DropperDeliver a second-stage payloadUsually none, hands off to the payload
Spyware / InfostealerCollect and exfiltrate dataLightweight, favors staying unnoticed

Treat this table as a reference for vocabulary, not a filing system for every sample you'll encounter. Real-world malware routinely blends categories rather than fitting one label cleanly. A commodity loader drops a backdoor that also functions as a rootkit-lite by hooking a handful of API calls to hide its own process. A ransomware operation deploys an infostealer days before detonating the encryptor, so the same intrusion produces artifacts belonging to two different classification categories at two different phases. Analysts who insist on forcing a single label onto a sample often miss capability the sample actually has, because they stop looking once the label feels satisfied. The more useful habit is describing behavior directly: what does this sample do, in what order, rather than what single word best describes it.

Building an Isolated Analysis Lab

Isolation is the single non-negotiable requirement before running any untrusted sample. The reason isn't abstract caution, it's containment.

Why isolation is non-negotiable:
  • A sample that reaches out to a real command-and-control server can alert the attacker that their malware has been discovered and is under analysis, changing their behavior and potentially burning an active investigation.
  • A sample with worm-like spreading behavior can move laterally into anything reachable on the same network.
  • A sample handled outside proper containment on infrastructure you don't fully control can create real legal and ethical exposure, since you're responsible for anything it touches once it starts running.

A workable lab needs a small number of core components, and none of them require exotic infrastructure.

Isolated Hardware or VM

The first component is an isolated virtual machine or a dedicated piece of hardware that never touches production systems or personal devices. Using your daily laptop, even with a VM on it, is a risk if that VM isn't properly isolated at the network layer, since a hypervisor escape or a misconfigured virtual network adapter is a much shorter path to compromise than most analysts assume.

Controlled Networking

The second component is controlled networking. The default posture for any analysis VM should be no real internet access at all. Malware that can't reach the internet still needs to think it can, though, because a lot of behavior only triggers once a network call succeeds or fails in a specific way. That's where tools like INetSim and FakeNet-NG come in: they simulate common internet services (DNS, HTTP, FTP, and others) on the isolated network, so a sample believes it's talking to the real internet and reveals its intended behavior, connection attempts, beacon patterns, requested URLs, without ever leaving the lab.

Snapshotting

The third component is snapshotting. Malware execution changes system state in ways that are tedious or impossible to fully undo by hand, from registry modifications to dropped files to services created and never cleaned up. Snapshotting turns that into a repeatable cycle instead of a one-time cleanup problem.

1
Snapshot
Take a clean snapshot of the VM before detonating anything.
→
2
Detonate
Run the sample inside the isolated VM.
→
3
Observe
Watch what it changes, drops, or contacts.
→
4
Revert
Restore the clean snapshot, discarding every change the run made, then repeat for the next sample.
Warning: Network isolation is necessary but not sufficient. Some malware actively checks whether it's running inside a virtual machine or a sandbox environment, and it changes or suppresses its malicious behavior when it detects one, precisely to defeat analysis like this. An isolated lab keeps a sample from causing damage, but it doesn't guarantee the sample will show you its real behavior. Chapter 8 covers anti-VM and anti-sandbox detection techniques, and the evasion arms race, in depth.

Static vs Dynamic Analysis at a Glance

Every technique in this module falls into one of two broad approaches, and understanding the difference now sets up the next two chapters directly.

Static AnalysisDynamic Analysis
ApproachExamines a file without executing itExecutes the sample in a controlled environment and observes what it does
What it coversHashes for identification, readable strings, file structure (the PE format, Chapter 2), disassembled code read for logicAPI calls, file system and registry changes, network traffic captured while the sample runs inside the isolated lab
StrengthFast, safe by default since nothing executes, produces a lot of useful information quickly on a straightforward sampleShows real behavior instead of inferred behavior, which matters when a sample's logic is hard to follow from code alone
Blind spotPacked or obfuscated code: if the malicious logic is compressed or encrypted until runtime, strings and disassembly tell you almost nothing, since what's visible on disk isn't what actually executesAnti-sandbox and anti-analysis tricks (checking for a debugger, checking for signs of virtualization, delaying activity past a sandbox's typical run time) can suppress the behavior you're trying to observe
What covers that blind spotDynamic analysis, because the code has to unpack itself into memory to run, revealing behavior static analysis alone can't seeStatic analysis of the same file's code, since the logic for detecting a sandbox has to exist somewhere in the binary
Tip: Don't think of static and dynamic analysis as two competing methods where you pick one. Think of them as two lenses you apply to the same sample, in whichever order makes sense for what you're looking at. Chapter 2 goes deep on static techniques, and Chapter 3 does the same for dynamic analysis and sandboxing, but most real investigations move back and forth between both throughout.

What This Module Covers

Eight chapters build malware analysis knowledge from the foundational concepts in this chapter through applied detection engineering. Chapter 1 (this chapter) established malware classification, lab isolation, the static versus dynamic distinction, and handling discipline. Everything after this builds directly on those four ideas.

  • Chapter 2, Static Analysis Fundamentals: hashing and fuzzy hashing, the PE file format, strings and import/export table analysis, and packer and signature detection tools.
  • Chapter 3, Dynamic Analysis and Sandboxing: sandbox tooling, process and API call monitoring, and network traffic capture.
  • Chapter 4, Unpacking and Deobfuscation: manual unpacking and original entry point identification, string and configuration deobfuscation, and where memory forensics picks up what disk-based analysis misses once code only exists unpacked in RAM.
  • Chapter 5, Malware Tradecraft: persistence mechanisms, process injection techniques from classic DLL injection through reflective loading, and the beaconing patterns behind command-and-control communication.
  • Chapter 6, Detection Engineering: turning analysis output into YARA rules, Sigma rules, and behavioral detection logic mapped to ATT&CK.
  • Chapter 7, Malware Threat Actor Campaigns: documented loader-to-ransomware ecosystems and commodity C2 frameworks repurposed across multiple actors.
  • Chapter 8, Advanced Evasion and Anti-Analysis: anti-VM and anti-debug techniques, AMSI and EDR evasion concepts, and fileless, living-off-the-land malware.
Tip: If you're coming from the Cloud Security or Threat Hunting modules, the analytical mindset transfers directly: form a hypothesis about what a sample does, gather evidence to test it, and revise as new evidence appears. What's different here is the evidence itself, binary structure, API calls, network behavior, rather than logs and identity signals. This module exists to teach you how to read that evidence.

Key Takeaways

  • Malware analysis serves three purposes: fast triage, deep-dive capability analysis, and producing IOCs and behavioral detail that feed detection engineering. Analysis that doesn't feed one of those outcomes isn't doing its job.
  • Malware types (virus, worm, trojan, ransomware, backdoor, rootkit, loader/dropper, spyware/infostealer) describe distinct strategies for harm, spread, and persistence, but real samples routinely blend multiple categories rather than fitting one label cleanly.
  • An isolated analysis lab needs an isolated VM or dedicated hardware, controlled networking (no real internet by default, simulated with tools like INetSim or FakeNet-NG), and snapshotting to revert cleanly after each run.
  • Network isolation is necessary but not sufficient. Some malware detects sandbox and VM environments and suppresses its own malicious behavior in response, a topic Chapter 8 covers in depth.
  • Static analysis (examining a file without running it) and dynamic analysis (executing it and observing behavior) each cover blind spots the other has, which is why the module dedicates separate chapters to both.
  • Sample handling follows enterprise IR discipline: documented chain of custody, never analyzing unknown samples on production or personal networks, conventional password-protected zip packaging, and authorization before analyzing anything not already confirmed malicious or provided for training.

Knowledge Check

Click an answer to reveal the explanation.

What is the clearest distinction between a worm and a virus?

A worm is self-propagating, it exploits a vulnerability or weak configuration to move from host to host without needing a user to run anything. A virus attaches to a legitimate host file and relies on that file being executed or shared to spread. That host-dependent versus self-propagating distinction is the classic dividing line between the two categories.

Why is network isolation alone not a guarantee that an analysis lab will reveal a sample's true behavior?

Network isolation contains the damage a sample could cause, but it doesn't stop malware from checking for signs of virtualization or sandboxing and altering or withholding its behavior when it detects them. This is a deliberate evasion strategy, and Chapter 8 covers the specific anti-VM and anti-sandbox techniques involved.

A sample is heavily packed, and its strings and disassembly are unreadable when examined statically. Which approach is most likely to reveal what it actually does?

Packed or obfuscated code defeats naive static analysis because the malicious logic isn't readable until it's unpacked into memory at runtime. Dynamic analysis executes the sample in the isolated lab and observes API calls, file and registry activity, and network behavior directly, which reveals the real behavior even when the on-disk code is unreadable.