Malware Analysis Foundations
Before you run a single tool against a real sample, you need a place to run it safely and a vocabulary for describing what it does once it runs. Most of the damage done during early malware analysis has nothing to do with missing a technical detail. It comes from analyzing a live sample on a machine that's still connected to something that matters. This chapter builds the foundation the rest of the module depends on: what malware analysis is actually for, how to classify what you're looking at, how to build an isolated lab you can detonate samples in without consequence, the difference between static and dynamic analysis, and the handling discipline that keeps the process safe and defensible.
What Is Malware Analysis, and Why Does It Matter?
Malware analysis is the process of examining a malicious file or a running malicious process to determine what it is, what it does, and what it's capable of doing to the environment it lands in. That sounds simple, but the phrase hides three separate jobs that analysts actually do, often on the same sample within the same hour.
| Job | Question It Answers | Pace and Method | Output |
|---|---|---|---|
| Triage | How bad is this, and how fast do I need to act? | Minutes, using surface-level static checks (Chapter 2) | A verdict: known commodity loader, targeted implant, or false positive |
| Deep-dive analysis | What is this sample's full capability? | Slower, deliberate, combines static and dynamic techniques | A complete behavioral picture: persistence, communication, data touched, evasion |
| Producing usable output | What can the rest of the security program act on? | Extracting indicators as the sample is worked, not as an afterthought | IOCs (hash, C2 domain, mutex name, persistence registry key) and behavioral detail for detection rules |
Triage: How Bad, How Fast
When a suspicious file lands on an analyst's desk, whether from an EDR alert, a user report, or an email attachment quarantined by a gateway, the first question isn't "how does this work." It's "how bad is this, and how fast do I need to act." A skilled analyst can often answer that in minutes using the surface-level static checks covered in Chapter 2.
Deep-Dive Analysis: Full Capability
Once a sample is confirmed malicious and worth the time, the goal shifts to understanding its full capability: what it persists as, what it communicates with, what data it touches, and how it evades detection. Static and dynamic techniques get combined here to build a complete picture of behavior rather than a quick verdict.
Producing Usable Output: The Real Product
Analysis that stays in an analyst's head helps nobody. The real product of malware analysis is indicators of compromise (IOCs) that other analysts can search for, and behavioral detail that a detection engineer can turn into a rule. A hash, a C2 domain, a mutex name, a registry key used for persistence: each of these, extracted correctly, becomes something a SOC can hunt for across an entire environment. Chapter 6 covers turning that output into YARA and Sigma rules directly, but the discipline of extracting clean, accurate indicators starts here, in how you approach the sample in the first place.
Malware Classification by Type
Classifying malware by type gives analysts a shared vocabulary and a starting set of expectations about behavior. Each category below describes a distinct strategy for causing harm, spreading, or maintaining access, though as the closing note in this section explains, real samples rarely respect these boundaries cleanly.
- Virus: attaches itself to a legitimate host file or program and relies on that host being executed or shared to spread. It typically needs some form of user action, running the infected program, opening the infected document, to propagate.
- Worm: self-propagating. It spreads across a network on its own, exploiting a vulnerability or a weak configuration to move from host to host without requiring a user to run anything. Host-dependent versus self-propagating is the classic line drawn between virus and worm, even though both terms get used loosely in casual conversation.
- Trojan: disguises itself as something legitimate or desirable to trick a user into running it voluntarily. Unlike a virus, it doesn't attach to a host file, and unlike a worm, it doesn't spread on its own. It depends entirely on social engineering to get executed the first time.
- Ransomware: encrypts a victim's files and demands payment for the decryption key. Modern ransomware operations frequently add data theft and extortion on top of encryption.
- Backdoor: gives an attacker persistent remote access to a compromised system, often installed as a secondary payload after initial access.
- Rootkit: focuses specifically on hiding, modifying the operating system's own reporting mechanisms so that its files, processes, or network connections don't show up in normal system views.
- Loader / dropper: whose entire job is to fetch or unpack a second-stage payload and execute it, functioning as the delivery mechanism rather than the final capability.
- Spyware / infostealer: focuses on collecting information, credentials, browser data, cryptocurrency wallet files, and exfiltrating it, usually without any destructive action that would tip off the victim.
| Type | Primary Goal | Typical Persistence |
|---|---|---|
| Virus | Spread via host file infection | Embedded in infected files |
| Worm | Self-propagate across a network | Often none beyond active spread |
| Trojan | Gain execution via deception | Varies by payload delivered |
| Ransomware | Extort via encryption and data theft | Short-lived, runs to completion |
| Backdoor | Maintain remote access | Registry run keys, scheduled tasks, services |
| Rootkit | Hide presence from the OS | Kernel or boot-level hooks |
| Loader / Dropper | Deliver a second-stage payload | Usually none, hands off to the payload |
| Spyware / Infostealer | Collect and exfiltrate data | Lightweight, favors staying unnoticed |
Treat this table as a reference for vocabulary, not a filing system for every sample you'll encounter. Real-world malware routinely blends categories rather than fitting one label cleanly. A commodity loader drops a backdoor that also functions as a rootkit-lite by hooking a handful of API calls to hide its own process. A ransomware operation deploys an infostealer days before detonating the encryptor, so the same intrusion produces artifacts belonging to two different classification categories at two different phases. Analysts who insist on forcing a single label onto a sample often miss capability the sample actually has, because they stop looking once the label feels satisfied. The more useful habit is describing behavior directly: what does this sample do, in what order, rather than what single word best describes it.
Building an Isolated Analysis Lab
Isolation is the single non-negotiable requirement before running any untrusted sample. The reason isn't abstract caution, it's containment.
- A sample that reaches out to a real command-and-control server can alert the attacker that their malware has been discovered and is under analysis, changing their behavior and potentially burning an active investigation.
- A sample with worm-like spreading behavior can move laterally into anything reachable on the same network.
- A sample handled outside proper containment on infrastructure you don't fully control can create real legal and ethical exposure, since you're responsible for anything it touches once it starts running.
A workable lab needs a small number of core components, and none of them require exotic infrastructure.
Isolated Hardware or VM
The first component is an isolated virtual machine or a dedicated piece of hardware that never touches production systems or personal devices. Using your daily laptop, even with a VM on it, is a risk if that VM isn't properly isolated at the network layer, since a hypervisor escape or a misconfigured virtual network adapter is a much shorter path to compromise than most analysts assume.
Controlled Networking
The second component is controlled networking. The default posture for any analysis VM should be no real internet access at all. Malware that can't reach the internet still needs to think it can, though, because a lot of behavior only triggers once a network call succeeds or fails in a specific way. That's where tools like INetSim and FakeNet-NG come in: they simulate common internet services (DNS, HTTP, FTP, and others) on the isolated network, so a sample believes it's talking to the real internet and reveals its intended behavior, connection attempts, beacon patterns, requested URLs, without ever leaving the lab.
Snapshotting
The third component is snapshotting. Malware execution changes system state in ways that are tedious or impossible to fully undo by hand, from registry modifications to dropped files to services created and never cleaned up. Snapshotting turns that into a repeatable cycle instead of a one-time cleanup problem.
Static vs Dynamic Analysis at a Glance
Every technique in this module falls into one of two broad approaches, and understanding the difference now sets up the next two chapters directly.
| Static Analysis | Dynamic Analysis | |
|---|---|---|
| Approach | Examines a file without executing it | Executes the sample in a controlled environment and observes what it does |
| What it covers | Hashes for identification, readable strings, file structure (the PE format, Chapter 2), disassembled code read for logic | API calls, file system and registry changes, network traffic captured while the sample runs inside the isolated lab |
| Strength | Fast, safe by default since nothing executes, produces a lot of useful information quickly on a straightforward sample | Shows real behavior instead of inferred behavior, which matters when a sample's logic is hard to follow from code alone |
| Blind spot | Packed or obfuscated code: if the malicious logic is compressed or encrypted until runtime, strings and disassembly tell you almost nothing, since what's visible on disk isn't what actually executes | Anti-sandbox and anti-analysis tricks (checking for a debugger, checking for signs of virtualization, delaying activity past a sandbox's typical run time) can suppress the behavior you're trying to observe |
| What covers that blind spot | Dynamic analysis, because the code has to unpack itself into memory to run, revealing behavior static analysis alone can't see | Static analysis of the same file's code, since the logic for detecting a sandbox has to exist somewhere in the binary |
Legal, Ethical, and Safety Handling Basics
Technical isolation solves the containment problem. It doesn't solve the handling problem, and enterprise incident response treats malware samples with the same discipline as any other piece of evidence.
Chain of Custody
In an IR context, samples typically move through a documented chain of custody: who collected it, from where, when, and who has handled it since. That record matters if the incident ever becomes the subject of a legal action, a regulatory inquiry, or an internal review, and getting into the habit of documenting it even in training exercises builds the right muscle memory for when it counts.
Never Analyze Outside the Lab
The single most important safety rule is one that's easy to state and surprisingly easy to violate under time pressure: never analyze an unknown or confirmed-malicious sample on a production network or a personal device. That includes plugging in a USB drive that might contain a sample, opening a suspicious attachment "just to look," or running a quick check on a work laptop that's still connected to the corporate network. Every one of those actions defeats the entire point of building an isolated lab in the first place.
Password-Protected Packaging
When samples need to move between people or systems, they're conventionally packaged in a password-protected zip archive, most commonly using the password "infected." This is a widely used industry convention, not a security control. Its actual purpose is to stop email gateways, endpoint AV, and file-sharing platforms from automatically detecting and deleting the sample in transit, which would otherwise make it impossible to share known-malicious files for legitimate analysis or training purposes.
Authorization Before Analysis
Underneath all of this sits a general principle worth internalizing early: get authorization before analyzing anything that isn't already confirmed malicious or explicitly provided for training. Running unauthorized executables against systems you don't have clear permission to test, even out of curiosity, can cross legal lines depending on jurisdiction and context. The samples used throughout this module, and any sample an analyst works within a professional setting, should come through an authorized, documented path: an IR engagement, a malware repository built for research, or training material provided for exactly this purpose.
What This Module Covers
Eight chapters build malware analysis knowledge from the foundational concepts in this chapter through applied detection engineering. Chapter 1 (this chapter) established malware classification, lab isolation, the static versus dynamic distinction, and handling discipline. Everything after this builds directly on those four ideas.
- Chapter 2, Static Analysis Fundamentals: hashing and fuzzy hashing, the PE file format, strings and import/export table analysis, and packer and signature detection tools.
- Chapter 3, Dynamic Analysis and Sandboxing: sandbox tooling, process and API call monitoring, and network traffic capture.
- Chapter 4, Unpacking and Deobfuscation: manual unpacking and original entry point identification, string and configuration deobfuscation, and where memory forensics picks up what disk-based analysis misses once code only exists unpacked in RAM.
- Chapter 5, Malware Tradecraft: persistence mechanisms, process injection techniques from classic DLL injection through reflective loading, and the beaconing patterns behind command-and-control communication.
- Chapter 6, Detection Engineering: turning analysis output into YARA rules, Sigma rules, and behavioral detection logic mapped to ATT&CK.
- Chapter 7, Malware Threat Actor Campaigns: documented loader-to-ransomware ecosystems and commodity C2 frameworks repurposed across multiple actors.
- Chapter 8, Advanced Evasion and Anti-Analysis: anti-VM and anti-debug techniques, AMSI and EDR evasion concepts, and fileless, living-off-the-land malware.
Key Takeaways
- Malware analysis serves three purposes: fast triage, deep-dive capability analysis, and producing IOCs and behavioral detail that feed detection engineering. Analysis that doesn't feed one of those outcomes isn't doing its job.
- Malware types (virus, worm, trojan, ransomware, backdoor, rootkit, loader/dropper, spyware/infostealer) describe distinct strategies for harm, spread, and persistence, but real samples routinely blend multiple categories rather than fitting one label cleanly.
- An isolated analysis lab needs an isolated VM or dedicated hardware, controlled networking (no real internet by default, simulated with tools like INetSim or FakeNet-NG), and snapshotting to revert cleanly after each run.
- Network isolation is necessary but not sufficient. Some malware detects sandbox and VM environments and suppresses its own malicious behavior in response, a topic Chapter 8 covers in depth.
- Static analysis (examining a file without running it) and dynamic analysis (executing it and observing behavior) each cover blind spots the other has, which is why the module dedicates separate chapters to both.
- Sample handling follows enterprise IR discipline: documented chain of custody, never analyzing unknown samples on production or personal networks, conventional password-protected zip packaging, and authorization before analyzing anything not already confirmed malicious or provided for training.
Knowledge Check
Click an answer to reveal the explanation.
What is the clearest distinction between a worm and a virus?
Why is network isolation alone not a guarantee that an analysis lab will reveal a sample's true behavior?
A sample is heavily packed, and its strings and disassembly are unreadable when examined statically. Which approach is most likely to reveal what it actually does?