CHAPTER 02 40 MIN READ INTERMEDIATE

Static Analysis Fundamentals

Chapter 1 drew the line between static and dynamic analysis at a glance. This chapter puts static analysis to work in practice: hashing a sample for identification, reading the PE file format well enough to know what a section table and an import table are actually telling you, pulling readable strings out of a binary, recognizing when a packer has made most of that inspection pointless, and disassembling the handful of functions that still need a direct read once everything else has narrowed the search. None of it requires running the sample, which is exactly the point. Static analysis is the fast, safe, repeatable first pass every other technique in this module builds on, and knowing its ceiling is as important as knowing its techniques.

PE format hashing strings analysis packer detection disassembly

Why Static Analysis Comes First

Static analysis means examining a sample without executing it. You read its bytes, its structure, and whatever text or metadata it carries, and you draw conclusions from what is sitting on disk rather than from what the sample does when it runs. That distinction is the entire reason static analysis is the first thing an analyst does with a new sample rather than the last.

Triage Economics

A sandbox run takes minutes, sometimes longer if a sample stalls waiting for a command-and-control callback that never arrives, and dynamic analysis infrastructure is a shared, finite resource in most SOCs. Static analysis takes seconds. Hashing a file, pulling its strings, and checking its import table can be done on hundreds of samples in the time a single sandbox detonation takes, which means static analysis is what decides which samples are worth spending that sandbox time on in the first place. A sample that resolves to a known hash already in your threat intel platform, with a known family and known behavior, usually does not need a fresh dynamic run at all. One that is unfamiliar, unpacked, and calling process-injection APIs by name just earned itself a spot in the queue.

The Safety Argument

Static analysis never runs attacker-controlled code on anything, so it carries none of the containment risk that comes with detonating a live sample, even inside an isolated lab. You can perform static analysis on a locked-down analysis workstation with no network access at all and lose nothing in the process.

The Honest Limitation

Static analysis only tells you what a file looks like, not what it does, and a sample built with packing or obfuscation in mind can make "what it looks like" almost meaningless. A packed executable's visible strings are mostly the packer's own runtime stub, not the payload's; its import table lists only the handful of APIs the unpacking stub needs, not the APIs the actual malware calls once it is running in memory. Static analysis against a well-packed sample can still tell you something, namely that the sample is packed, but it will not hand you the payload's real behavior. That is precisely the gap dynamic analysis exists to close, and it's the reason this chapter and the next one are a pair: static analysis tells you what to look for and where the ceiling is, dynamic analysis picks up from there.

Note: "Static first" is a triage default, not a rule that dynamic analysis is optional. A sample that looks clean statically because it is packed, encrypted, or simply well-obfuscated is often the sample that most needs a sandbox run, not the one that needs it least.

Hashing for Identification and Clustering

Hashing is usually the very first action taken on any sample, and it serves two distinct purposes that are worth separating clearly: exact identification and similarity clustering.

Cryptographic Hashes: Exact Identification

Cryptographic hashes, MD5, SHA1, and SHA256, take a file's exact bytes and produce a fixed-length digest. Change a single byte anywhere in the file and the digest changes completely. That property makes cryptographic hashing perfect for exact identification: if a hash you compute matches a hash in VirusTotal, your threat intel platform, or an internal case file, you are looking at the identical file someone else already analyzed, byte for byte. MD5 and SHA1 are still widely used as lookup keys purely out of legacy convention, most threat intel platforms and OSINT sources index by all three, but SHA256 is the stronger choice when you have to pick one, since MD5 and SHA1 both have known collision weaknesses that matter for integrity guarantees even if they rarely matter for day-to-day malware lookups. Whichever algorithm you use, the hash is only useful if the exact file has been seen and catalogued before. Change one byte, recompile with a different string, or run the file through a different packer, and the hash tells you nothing, even if the underlying malware is functionally identical.

Ssdeep: Fuzzy Hashing for Similarity

That brittleness is exactly what fuzzy and similarity hashing is built to work around. Ssdeep computes a hash based on a file's content in a way that produces similar output for similar files, so two samples that share most of their bytes but differ in a few places will show a high similarity score even though their cryptographic hashes are completely unrelated. It's a useful tool for spotting minor variants of a known sample, a recompiled build, a slightly patched configuration, a dropper with one string changed, that a cryptographic hash would treat as entirely unrelated files.

Imphash: Clustering by Import Table

Imphash works on a narrower but often more useful slice of the file: the import table specifically. It hashes the ordered list of API names (and the DLLs they come from) that a PE file imports, so two files with completely different code, different strings, and different cryptographic hashes will produce the same imphash if they were built with the same compiler, linker, and library set in the same order, which is common when malware is produced by the same builder, packer, or toolkit. That makes imphash a strong clustering signal for attribution: analysts use it to group samples that came out of the same builder even when the payload itself was rebuilt or repacked between samples. It's not infallible, legitimate software compiled with common toolchains can share an imphash purely by coincidence, and a sample that statically links its imports or resolves APIs dynamically at runtime (rather than importing them directly) can produce a meaningless or empty imphash. But when it lines up across a set of samples that share other indicators too, it's one of the more durable clustering signals static analysis offers.

Hash TypeWhat It MeasuresBest Used For
MD5 / SHA1 / SHA256Exact byte-for-byte content of the whole fileConfirming you have the identical file someone else already analyzed; VirusTotal/OSINT lookups
ssdeepContent similarity across the whole file (fuzzy hash)Spotting minor variants, recompiles, or patched configs of a known sample
imphashOrdered list of imported APIs and their source DLLsClustering samples built with the same builder/toolkit, even with different payloads
Tip: Compute all three when you first receive a sample and record them together in your case notes. A cryptographic hash miss followed by an imphash or ssdeep hit is often the first clue that you're looking at a new variant of a family you already have coverage for, rather than something genuinely novel.

The PE File Format

Most Windows malware ships as a Portable Executable, the same file format used by every legitimate .exe and .dll on the operating system, and reading a PE file's structure is one of the highest-value skills in static analysis because so much of what a sample can do is declared in that structure before a single instruction runs.

DOS Header

A PE file opens with a DOS header, a holdover from MS-DOS compatibility that every valid PE file still carries. Its first two bytes are the ASCII characters "MZ", the initials of Mark Zbikowski, one of the format's original designers, and that MZ signature is the fastest sanity check available: if a file claiming to be an executable doesn't start with MZ, either it isn't one or something has altered it. The DOS header also contains a pointer, e_lfanew, to where the real PE header begins further into the file, since the DOS header itself exists mostly to keep old DOS systems able to print a friendly error if they tried to run a Windows program directly.

PE Header and the Compile Timestamp

The PE header proper begins with its own signature, the ASCII characters "PE" followed by two null bytes, and from there declares the file's machine type (x86, x64, ARM), the number of sections that follow, and a timestamp field that records when the file was compiled. That timestamp is worth flagging early and clearly: it is trivial to forge. Malware authors routinely zero it out, set it to a nonsense far-future or far-past date, or backdate it to match a legitimate vendor's build window specifically to throw off an analyst's timeline. Treat the PE timestamp as a data point worth noting, never as ground truth for when a sample was actually built.

Section Table

Below the header sits the section table, which lists each section the file is divided into along with its name, its size, and the memory permissions it will be mapped with once loaded. Section names are a convention, not an enforced rule, so a sample can name a section anything it wants, but the common ones are worth recognizing on sight because a mismatch between a section's name and its actual content or permissions is itself a signal.

SectionTypical Purpose
.textExecutable machine code (typically read + execute)
.dataInitialized global and static variables (read + write)
.rdataRead-only data: string constants, the import table, debug info
.rsrcResources: icons, version info, embedded dialogs, sometimes embedded payloads
.relocBase relocation table used when the file loads at a non-preferred base address
.bssUninitialized data, reserved but not stored in the file itself

Import Table, Export Table, and Entry Point

Two structures inside the PE file matter more than any other for a static analyst: the import table and the export table. The import table lists every external function the file calls and which DLL each one comes from, which means it is, in effect, a declared list of what the sample is capable of doing before it ever runs. A file importing WinInet or WinHTTP functions has network capability. A file importing VirtualAlloc, WriteProcessMemory, and CreateRemoteThread together has the building blocks for process injection. The export table, by contrast, lists functions the file makes available to other code, relevant mainly for DLLs, and in malware it sometimes shows up as a single oddly-named export used by a loader to call into the payload, or is absent entirely when the file is a straightforward standalone executable. Finally, the entry point, recorded in the PE header as a relative virtual address, marks exactly where execution begins once the loader hands off control, and comparing that address against where you'd expect a normal compiled program's entry point to sit is one of the quickest ways to notice a packer stub has been placed in front of the real code.

Warning: Never treat the PE compile timestamp as reliable evidence for a campaign timeline or attribution argument on its own. It is one unauthenticated field in a file the adversary fully controls, and forging it costs them nothing.

Strings Analysis

Running a strings extraction against a binary, whether with the classic Unix strings utility, its Windows-native equivalents, or a GUI tool that does the same thing, pulls out every sequence of printable characters embedded in the file. For an unpacked, unobfuscated sample this can be extraordinarily informative on its own, often faster than any other single step in the whole triage process.

What Strings Reveal

  • Hardcoded URLs and IP addresses can point straight at command-and-control infrastructure.
  • File paths reveal where a sample expects to drop additional components or where it was built on the developer's own machine, sometimes leaking a username or project directory in the process.
  • Registry key paths hint at intended persistence locations before you've run the sample once.
  • Windows API error message text and format strings, pulled in as a side effect of static linking, hint at what a sample was built to attempt even before you trace its imports directly.
  • Mutex names and configuration strings show up in plain text every so often simply because the author didn't bother to hide them.

The Limitation

Strings extraction is only as good as what's actually stored as plain, uncompressed, unencrypted text in the file. Run it against a packed or encrypted sample and you get mostly noise, fragments of the packer's own stub, compiler and linker artifacts, and none of the payload's real strings, because the payload itself is sitting compressed or encrypted in the file and doesn't exist as readable text until it's unpacked in memory at runtime.

Stack Strings and FLOSS

Modern malware compounds the problem even without full packing, by building strings at runtime instead of storing them directly: pushing individual characters or bytes onto the stack and assembling them in memory just before use, a technique generally called stack strings. Plain strings extraction is blind to this, since there's no contiguous readable string sitting in the file to find, only a sequence of instructions that reconstructs one. FLOSS (FLARE Obfuscated String Solver), built by Mandiant's FLARE team, was built specifically to close that gap: it emulates the relevant code paths well enough to reconstruct stack strings, resolve simple string-decoding routines, and decode strings hidden behind common obfuscation patterns that a plain strings pass would never surface, all without executing the sample for real.

TEXT
$ strings -n 6 sample.exe | grep -Ei "http|\.dll|Software\\\\"

http://185.203.xx.xx/gate.php
KERNEL32.dll
ADVAPI32.dll
WININET.dll
Software\Microsoft\Windows\CurrentVersion\Run
%s: connection failed, retrying in %d seconds
MicrosoftEdgeUpdateTaskMachine
GetProcAddress
LoadLibraryA

That kind of output, illustrative rather than pulled from any specific real family, is exactly the pattern worth reading closely: a plausible C2 URL, a persistence-relevant registry path, a scheduled-task-style name likely used for masquerading, and a pair of imports (LoadLibraryA and GetProcAddress) that together suggest the sample resolves at least some of its API calls dynamically rather than importing them directly, which is itself worth remembering when you get to the import table and notice it looks shorter than expected.

Tip: When plain strings output looks unusually short or unusually clean for the file's size, that's a signal in itself. A five-hundred-kilobyte binary that yields twenty strings, mostly compiler boilerplate, is very likely packed, and the right next step is entropy or packer detection, not a longer strings command.

Import and Export Tables as a Behavioral Preview

If strings analysis gives you a sample's roadmap, the import table gives you its declared capability list, and reading it carefully is one of the most efficient ways to form an early hypothesis about what a sample is built to do, well before any dynamic analysis confirms it.

Red-Flag API Combinations

Certain API combinations are recognized red flags in the analyst community precisely because they appear together so consistently in known malicious tradecraft. VirtualAlloc, WriteProcessMemory, and CreateRemoteThread appearing together is the classic process-injection triad: allocate memory in a target process, write code or data into it, then create a thread to execute it there. Seeing all three imported by the same sample is a strong reason to expect injection behavior once you get to dynamic analysis. InternetOpen, InternetConnect, and the broader WinINet or WinHTTP families signal straightforward network capability, worth correlating against whatever URLs or IPs showed up in the strings pass. CryptEncrypt, CryptAcquireContext, and the modern BCrypt equivalents signal cryptographic capability, which in the context of an unfamiliar unsigned binary is worth treating as a potential ransomware or data-staging indicator worth chasing further.

Imported API (or set)What It Suggests
VirtualAlloc / VirtualAllocExAllocating memory, often a precursor to code injection or unpacking a payload in memory
WriteProcessMemory + CreateRemoteThreadClassic process-injection pattern, especially when paired with VirtualAllocEx
SetWindowsHookExHooking, sometimes used for keylogging or injecting code into other processes' address space
InternetOpen / WinHttpOpen familyOutbound network capability; correlate against strings for C2 URLs/IPs
CryptEncrypt / BCryptEncrypt familyCryptographic capability; worth investigating for ransomware or data-staging behavior
RegSetValueEx / RegCreateKeyExRegistry write capability, often tied to persistence
LoadLibraryA + GetProcAddressDynamic API resolution at runtime, often used to hide the true import table from static tools

The Caveat: Context Over Single APIs

None of these APIs are exclusive to malware, and treating any single one as proof of malicious intent will produce false positives constantly. Legitimate software injects code for entirely benign reasons: application compatibility shims, some antivirus and EDR products themselves, debugging and profiling tools, and plenty of ordinary installers all call VirtualAllocEx and WriteProcessMemory as part of normal operation. CryptEncrypt shows up in any application that handles credentials or protects local data. The value of import table analysis is in context and combination, not in any one API name in isolation: the process-injection triad appearing together in a sample with no legitimate reason to inject code, alongside network imports and strings pointing at an unfamiliar external host, builds a coherent behavioral hypothesis. The same single API sitting alone in an otherwise unremarkable file usually doesn't.

Dynamic API Resolution as Its Own Signal

One more wrinkle worth flagging while it's fresh from the strings section above: a sample that imports only LoadLibraryA and GetProcAddress, with a suspiciously short import table otherwise, is very likely resolving its real capability list dynamically at runtime rather than declaring it statically. That's a deliberate evasion pattern precisely because it defeats import table analysis, and it's a strong signal on its own that dynamic analysis is going to carry more weight for this particular sample than static analysis will.

Packer and Signature Detection

What Packing Does

Packing compresses or encrypts a file's real payload and wraps it in a small stub that decompresses or decrypts that payload back into memory at runtime, then transfers execution to it. Crypters do the same conceptual job with a heavier emphasis on encryption specifically, often layered with additional anti-analysis tricks. Packing exists legitimately, commercial software has used it for decades to shrink distribution size and make casual reverse engineering harder, but malware authors adopted it just as enthusiastically for the second reason: a packed sample defeats exactly the static techniques covered earlier in this chapter. Strings analysis against a packed file mostly surfaces the packer's own runtime code, not the payload's. The import table shows only the handful of APIs the unpacking stub itself needs, typically little more than LoadLibraryA, GetProcAddress, and a memory allocation call, none of which reveal what the payload does once it's actually running.

Detecting that a file is packed, even without unpacking it yet, is itself a valuable static analysis outcome, because it tells you definitively that further static inspection has hit its ceiling and the next move needs to be either dynamic analysis or manual unpacking, the subject of Chapter 4. A few approaches are standard practice for making that call.

Entropy Analysis

Entropy analysis measures how random a section's byte content looks, on a scale that in practice runs from roughly 0 (highly ordered, repetitive data) to 8 (statistically indistinguishable from random noise). Compressed and encrypted data both produce high entropy, close to that upper bound, because compression and encryption are both designed to remove the redundancy and predictable patterns that make data compressible or analyzable in the first place. A section named .text with entropy sitting up near 7.9 or higher is a strong indicator that whatever is actually in that section isn't ordinary compiled machine code, machine code has enough structural repetition (common instruction sequences, aligned function prologues) to sit meaningfully lower, but is compressed or encrypted payload data instead.

Signature and Heuristic Detection

Signature-based detection takes a different approach, matching a file's entry point bytes or specific structural fingerprints against a database of known packer signatures. PEiD, though its original signature database hasn't kept pace with modern packers on its own, established the model that later tools built on: identify the specific packer or compiler a file was built with by matching known byte patterns at the entry point. Detect It Easy, generally known as DiE, is the actively maintained tool most analysts reach for today, combining signature matching with entropy visualization per section and heuristic detection for packers that don't match a known signature outright, giving a fast, fairly reliable first read on whether a file is packed and, when the signature hits, which packer produced it.

ApproachWhat It ChecksTypical Tooling
Entropy analysisRandomness of byte content per section; high entropy suggests compression or encryptionDetect It Easy (DiE), CFF Explorer, pestudio
Signature matchingKnown byte patterns at the entry point matching a cataloged packer or compilerPEiD-style databases, Detect It Easy
Heuristic detectionStructural anomalies (unusual section counts, mismatched section names/permissions, tiny import tables) not tied to a specific known signatureDetect It Easy, manual review

What a Packer Hit Means Next

None of these approaches unpack the file. What they give you is a confident answer to a narrower but still valuable question: is this sample packed, and if a signature matches, by what. That answer is a triage decision point, not an endpoint. A confirmed packer hit means static analysis on the payload itself is done for now, and the next step is either detonating the sample in a sandbox to observe what the unpacked payload does at runtime, or manually unpacking it to get the payload back onto disk for direct static inspection, which is exactly where Chapter 4 picks up.

Warning: High entropy in a resource section alone is a weak signal by itself. Legitimate files routinely embed compressed icons, images, or other pre-compressed resources in .rsrc, which will read as high entropy without the file being packed. Weigh entropy findings against which section shows it, .text and .data reading high is far more telling than .rsrc alone.

Disassembly and Decompilation Basics

Everything so far in this chapter tells you about a sample from the outside: what it's called, how it's shaped, what it says about itself in cleartext, whether it's hiding its real payload behind a packer. None of that reads the actual logic the sample executes. Disassembly does. A disassembler converts a binary's raw machine code back into assembly, the human-readable mnemonics (mov, call, jmp, cmp) that correspond one-to-one with what the CPU actually does instruction by instruction. It's slower and requires more background than anything covered above, but it's the only static technique that shows you the sample's decision logic directly rather than inferring it from surrounding evidence.

Tools: Ghidra, IDA Pro, Radare2

Ghidra, released and maintained by the NSA, and IDA Pro, the long-standing commercial standard, are the two disassemblers most analysts reach for. Both go a step further than raw disassembly with a built-in decompiler, which reconstructs pseudo-C from the assembly: instead of reading a page of mov and cmp instructions, you read something closer to if (result == 0) { call_function(); }. Decompiler output is never as exact as the real source and it occasionally gets a construct wrong, but for understanding what a function is trying to do, it's dramatically faster to read than raw assembly, which is why most triage-level reverse engineering leans on the decompiler view first and drops to raw assembly only when the pseudo-C looks wrong or incomplete. Radare2 and its Cutter GUI front-end offer a free, scriptable alternative with a smaller learning curve gap between the command line and a full GUI workflow.

Working from Cross-References, Not Top to Bottom

Reading a disassembly top to bottom, function by function, doesn't scale, and it isn't how this actually gets done in practice. The strings and import table work from earlier in this chapter exist specifically to give you a starting point: every disassembler lets you jump to the cross-references (xrefs) of a specific string or a specific imported API, landing you directly in the function that uses it. A suspicious string like a hardcoded C2 domain, or a sensitive import like CreateRemoteThread or VirtualAllocEx, becomes an anchor. Following its xrefs takes you straight to the handful of functions that actually matter, out of what might be hundreds or thousands of functions in the full binary, and that targeted approach is what makes static reverse engineering practical inside a triage timeline instead of a multi-day research exercise.

Tip: Full manual reverse engineering of an entire binary is rarely the goal, and it directly contradicts the triage economics this chapter opened with. Most disassembly work in a SOC or IR context is surgical: confirm what a handful of suspicious functions do, verify a decompiled routine matches the behavior dynamic analysis already showed you, or recover a hardcoded value like an XOR key or a C2 domain that string extraction alone couldn't surface because it's built at runtime. Treat disassembly as the tool you reach for when strings, imports, and dynamic behavior together still leave a specific question unanswered, not as a replacement for the faster techniques earlier in this chapter.

Bringing It Together: A Static Analysis Triage Workflow

Each technique in this chapter answers a narrow question on its own. Strung together in a consistent order, they form a repeatable first-pass workflow that takes a few minutes per sample and tells you, reliably, whether you're looking at something already known, something worth an immediate deeper look, or something static analysis has already told you as much as it's going to.

1
Hashing
Compute MD5, SHA1, SHA256, ssdeep, and imphash; check against VirusTotal and internal case history. A hit can end triage in seconds.
→
2
PE Structure Review
Confirm MZ/PE signatures, review the section table for name/content mismatches, note (but don't trust) the compile timestamp, and pull the import table.
→
3
Import/Export Analysis
Read the import table for the API patterns covered above, treating any hits as a hypothesis to carry forward.
→
4
Strings Extraction
Run strings, or FLOSS if the output looks thin, for URLs, paths, registry keys, and configuration values; cross-reference against imports.
→
5
Packer/Entropy Detection
Run this regardless of how earlier steps went. A packer hit reframes a short import table and thin strings as hidden behavior, not simplicity.

Where this workflow ends is exactly where Chapter 3 begins. A sample that comes back unpacked, with a readable import table and informative strings, has probably already told you most of what static analysis has to offer, and the remaining question is confirming behavior through a sandbox run. A sample that comes back packed has told you static analysis has hit its ceiling, and dynamic analysis, or manual unpacking, isn't optional anymore, it's the only way forward.

Key Takeaways

  • Static analysis examines a sample without running it, which makes it fast and safe, but packed, encrypted, or heavily obfuscated payloads limit how much it can actually reveal.
  • Cryptographic hashes (MD5/SHA1/SHA256) give exact identification for threat-intel lookups; ssdeep and imphash find related samples that differ in body but share code similarity or a common builder's import table.
  • The PE format's section table, import table, and entry point carry most of the analytic value; the compile timestamp is trivially forged and should never be treated as reliable evidence.
  • Strings analysis surfaces URLs, paths, registry keys, and configuration values in plaintext, but yields mostly noise against packed samples; FLOSS recovers stack strings and simple obfuscated strings that plain strings misses.
  • Import table APIs like the VirtualAlloc/WriteProcessMemory/CreateRemoteThread injection triad are worth investigating in context and combination, not treated as proof of malice on their own, since legitimate software calls the same APIs.
  • Packer and entropy detection (PEiD-style signatures, Detect It Easy, section entropy) tells you when static analysis has hit its ceiling and dynamic analysis or manual unpacking, covered in Chapters 3 and 4, needs to take over.

Knowledge Check

Click an answer to reveal the explanation.

Two malware samples have completely different SHA256 hashes and different code bodies, but the same imphash. What does that most likely indicate?

Imphash hashes the ordered list of imported APIs and their source DLLs, not the file's content or code body. A shared imphash across samples with otherwise unrelated hashes and code is a strong clustering signal that they came from the same builder, packer, or toolkit, which is exactly why analysts use it for attribution even when the payload itself was rebuilt or repacked between samples.

A sample's import table shows only LoadLibraryA and GetProcAddress, with almost nothing else, and its strings output is unusually short for a file of its size. What should an analyst conclude from this pattern?

LoadLibraryA and GetProcAddress are the two APIs a packer stub or an evasive sample needs to resolve its real capability list dynamically at runtime instead of declaring it in the static import table. Combined with unusually thin strings output for the file's size, this is a strong signal the sample is packed or deliberately hiding its behavior, and the next step should be entropy/packer detection followed by dynamic analysis or manual unpacking, not a conclusion that the file is simple.

Why is a PE file's compile timestamp considered unreliable evidence during static analysis?

The PE header's timestamp field is an unauthenticated value the file's author fully controls, and forging it costs nothing. Malware authors zero it out, set it to nonsense dates, or backdate it to match a legitimate vendor's build window specifically to mislead an analyst's timeline. It's worth recording as a data point, but should never be treated as ground truth for when a sample was actually compiled.