Malware and Campaign Intelligence
Malware intelligence is one of the most tactically useful CTI disciplines because it directly produces detection artifacts: YARA rules, behavioral indicators, and sandbox-extracted IOCs. This chapter covers how to classify malware families, read sandbox reports for intelligence value, cluster samples to campaigns, and write basic YARA rules without being a reverse engineer.
Malware Family Classification
Malware taxonomy matters for CTI because the family classification tells you what the malware is designed to do, what actor categories typically use it, and what behavioral indicators to look for in your telemetry. Knowing that a sample is a loader is different from knowing it is a RAT, which is different from knowing it is ransomware, even if you have not analyzed the binary itself yet.
| Family | Primary Function | Examples | Notes |
|---|---|---|---|
| Loader | Downloads and executes a second-stage payload | Emotet, IcedID, BazarLoader, Qakbot | Often delivered via phishing; small and designed to evade initial detection. Identifying a loader means a second-stage payload likely exists or was attempted, and the hunt scope should include what the loader retrieved. |
| Remote Access Trojan (RAT) | Persistent remote control of a compromised system | AsyncRAT, njRAT, Remcos, Agent Tesla | Typically includes keylogging, screenshot capture, file system access, command execution, and C2 communication. Commercial RATs are frequently used by lower-sophistication actors; advanced actors tend to use custom implants or frameworks like Cobalt Strike. |
| Backdoor | Maintains a persistent shell or command execution channel | Many custom nation-state implants | More minimal than a RAT, providing a persistent foothold while minimizing detection footprint. The distinction between RAT and backdoor is sometimes blurred in vendor reporting. |
| Infostealer | Credential and data theft | RedLine, Raccoon, Vidar, Lumma Stealer | Targets browser-stored passwords, cookies, credit card data, cryptocurrency wallets, FTP credentials, and email account data. Commonly sold as malware-as-a-service. Stolen logs are a significant initial access vector for subsequent intrusions. |
| Ransomware | Encrypts files and demands payment | LockBit, BlackCat/ALPHV, Cl0p, Play, Black Basta | Modern operations typically involve double extortion (encrypt and threaten to publish stolen data). Heavily organized around Ransomware-as-a-Service (RaaS), where a developer licenses the platform to affiliates. Pre-encryption activity follows predictable, huntable patterns. |
| Wiper | Destroys data rather than encrypting it for ransom | NotPetya, WhisperGate, HermeticWiper | Primarily used by nation-state actors in destructive operations. Designed to cause maximum damage quickly rather than for stealth; presence is typically discovered during or immediately after the destructive phase. |
Reading Sandbox Reports
Sandbox reports are automated behavioral analysis outputs produced by running a malware sample in an instrumented virtual environment and observing what it does. They are valuable even for analysts without reverse engineering skills because they surface behavioral indicators, network IOCs, persistence mechanisms, and file system changes without requiring manual code analysis. Most public sandbox platforms (Hybrid Analysis, Any.run, Triage, Joe Sandbox) produce rich reports that a CTI analyst can extract significant value from in under 20 minutes.
Process Tree Analysis
The first place to look. The process tree shows parent-child process relationships from execution. Suspicious patterns include: document process (Word, Excel, PDF reader) spawning a scripting engine (PowerShell, cmd.exe, wscript.exe), followed by that scripting engine spawning a download tool (certutil, curl, mshta) or injecting into a system process. Each step in an unusual process tree is a potential detection point and maps to a specific ATT&CK technique.
Network Indicators
C2 domains and IPs, HTTP request paths and User-Agent strings, DNS queries made during execution, and any certificates or protocol fingerprints captured in the simulated traffic. C2 domains observed in a sandbox report are high-confidence malicious indicators: the malware reached out to them during execution. Note the exact HTTP paths, User-Agent strings, and connection patterns: these are more durable than the domain itself if the actor rotates infrastructure.
Persistence Mechanisms
Documented in sandbox reports, these tell you exactly how to hunt for the malware in an already-compromised environment. Common persistence indicators include: registry Run key modifications (HKCU\Software\Microsoft\Windows\CurrentVersion\Run), scheduled task creation, service installation (Windows Services), startup folder file drops, and COM object hijacking. Each of these maps to Windows Event IDs or EDR events that can be searched retroactively.
File System Changes
Dropped files, created directories, and modified system files. The paths where malware drops secondary payloads are often consistent across samples within a family: a family that consistently drops to C:\Users\Public\ or C:\ProgramData\Microsoft\ is providing a high-signal hunt location. File naming conventions for dropped payloads also often persist within a family.
| Sandbox Report Section | Intelligence Value | Detection Application |
|---|---|---|
| Process tree | Execution chain, LOL technique usage | Parent-child process detection rules |
| Network connections | C2 IPs, domains, User-Agents, URI paths | Network IOCs, protocol detection, User-Agent rules |
| Registry modifications | Persistence mechanism type and key path | Registry monitoring rules, EventID 13 (Sysmon) |
| File drops | Staging paths, secondary payload names | File path monitoring, file creation rules |
| API calls | Code behavior, evasion techniques | EDR behavioral detection, YARA rule string candidates |
Campaign Clustering
Campaign clustering is the process of grouping related malware samples, infrastructure, and incidents into a coherent campaign picture. Individual samples are data points. Clustering them reveals patterns: this family of samples all use the same packer, share C2 infrastructure, were distributed in the same phishing wave, and target the same sector. The cluster is a campaign; the campaign belongs to an actor; the actor has a profile that predicts what they will do next.
| Clustering Method | Groups Samples By | Primary Tools | Notes |
|---|---|---|---|
| Code similarity | Shared significant code regions: a function, a crypter, a string decryption routine, or a persistence mechanism | BinDiff, YARA, Malpedia's code similarity feature | High similarity in non-trivial code regions is strong evidence of a common developer or variant family. More durable than hash clustering because it survives recompilation. |
| Infrastructure | Shared C2 infrastructure: the same IP, domain, or hosting provider | Passive DNS, certificate analysis | Can link samples that share no code, because the attacker deployed multiple distinct tools but pointed them at the same backend infrastructure. |
| Behavioral | What samples do rather than what they are: the same persistence mechanism, injection technique, anti-analysis technique, and exfiltration method | YARA rules targeting behavioral patterns rather than specific strings | Clusters samples together even if their code differs significantly, at scale. |
Attribution From Clustering Is Probabilistic
Two samples sharing a rare crypter and the same C2 infrastructure from the same hosting provider are likely linked. Whether they are linked to a specific named actor requires additional evidence. Good clustering practice produces activity clusters first, then assesses whether those clusters match known actor profiles, rather than starting with attribution and working backward.
YARA Rule Basics
YARA is a pattern-matching tool designed for malware identification and classification. A YARA rule defines conditions that must be true for a file to be considered a match. Rules can match on byte sequences, text strings, regular expressions, file structure characteristics (PE header fields, section names), or combinations of these. You do not need to be a reverse engineer to write useful YARA rules: many effective rules can be written directly from sandbox reports and behavioral analysis without touching a disassembler.
The Three Sections of a Rule
- meta: descriptive information, author, date, description, and any other metadata fields the rule author wants to include.
- strings: defines named string variables, the patterns to search for.
- condition: the Boolean logic specifying which combinations of strings must be present for the rule to fire.
rule Malware_SampleFamily_PersistenceKey {
meta:
author = "H3AD-SEC"
date = "2026-06-01"
description = "Detects SampleFamily malware by registry persistence string and C2 user-agent"
reference = "https://vendor-report-url"
confidence = "high"
strings:
$reg_key = "SOFTWARE\\Microsoft\\Windows\\CurrentVersion\\Run\\SvcHostUpdate" nocase
$ua_string = "Mozilla/5.0 (compatible; SampleBot/1.0)" nocase
$mutex = "Global\\SAMPLEFAM_MUTEX_01" wide
condition:
uint16(0) == 0x5A4D // Must be a PE file (MZ header)
and (2 of ($reg_key, $ua_string, $mutex))
}String Extraction From Sandbox Reports
The fastest route to YARA rule candidates. Look for:
- Mutex names (often unique per malware family)
- Registry key paths used for persistence
- User-Agent strings used in C2 communication
- URL paths in network connections
- Hardcoded strings in API call parameters
Strings that are unusual, specific, and unlikely to appear in legitimate software are the strongest YARA candidates. Strings that are generic (like "Windows" or "Mozilla") will generate massive false positive rates and should be avoided or combined with more specific conditions.
Context-Based Rules Reduce False Positives
Context-based rules require fewer strings to be useful because they anchor string matches to their position within the file. The uint16(0) == 0x5A4D condition (PE header check) ensures the rule only fires on Windows executables. Conditions like $string at entrypoint or $string in (0..1024) anchor matches to specific file regions. These context conditions dramatically reduce false positives without sacrificing detection coverage.
Malware Intelligence Feeds
Several public and commercial platforms provide ongoing malware intelligence that is useful without requiring in-house analysis capabilities. Understanding what each platform provides and how to consume it effectively is part of operating an intel-driven CTI function.
| Platform | Type | What It Provides |
|---|---|---|
| MalwareBazaar (abuse.ch) | Free community platform | Malware samples and associated indicators, searchable by hash, tag, family name, or signature, with JSON API access. Particularly useful for newly observed samples, often within hours of initial discovery, before vendor analysis is complete. |
| Malpedia (Fraunhofer FKIE) | Curated cross-vendor reference | A library of malware families with actor associations, YARA rules, and references to analysis reports. Each family page aggregates vendor reporting, lists known aliases across vendors, maps to ATT&CK Groups, and provides reference samples. The authoritative cross-vendor malware family reference, equivalent to ATT&CK Groups but for malware. |
| VirusTotal Intelligence | Paid tier | Full VirusTotal dataset access including YARA hunting (search for files matching a YARA rule across submissions), retrohunting (apply a YARA rule against historical submissions), and similarity search (find files similar to a seed sample). Enables proactive malware discovery rather than reactive indicator matching. |
| OALabs, Any.run, Triage | Analysis and sandbox platforms | OALabs (Open Analysis Live) publishes detailed malware analysis reports with extracted configuration data, C2 indicators, and YARA rules, with particularly strong coverage of commodity RATs and infostealers; its UnpacMe service automates unpacking of common packers. Any.run and Triage provide interactive sandbox environments for observing malware execution in real time. |
Integrate Feeds Into Your Workflow
Consuming malware intelligence feeds effectively requires integration with your analysis workflow. A YARA hunting feed from MalwareBazaar integrated with an EDR platform allows new malware families to be automatically searched across endpoint telemetry without manual effort. A Malpedia family update integrated with an actor profile can automatically flag when a previously profiled actor deploys a new tool variant. These integrations transform passive feed consumption into active threat detection.
Key Takeaways
- Malware family classification: loader (delivers second stage), RAT (remote control), backdoor (persistent shell), infostealer (credential/data theft), ransomware (encrypt for payment), wiper (destructive). Each family type implies different actor categories and response priorities.
- Sandbox reports produce behavioral IOCs without requiring reverse engineering: process trees, network connections, registry modifications, file drops, and API calls all provide detection-ready intelligence.
- Campaign clustering links samples to a common actor via code similarity, shared infrastructure, or behavioral overlap. Clustering produces activity groups before attribution, which keeps the analysis honest.
- YARA rules can be written from sandbox report strings: mutex names, registry paths, User-Agent strings, and URL patterns are strong candidates. Anchor string conditions with PE header checks and file offset constraints to reduce false positives.
- Key malware intelligence platforms: MalwareBazaar (free samples and indicators), Malpedia (family reference with actor mappings), VirusTotal Intelligence (YARA hunting), OALabs and Any.run (detailed behavioral analysis).
- Integrate malware feeds programmatically where possible: YARA hunting against live endpoint telemetry converts passive feed consumption into active detection.
Knowledge Check
Click an answer to reveal the explanation.
A sandbox report for a new sample shows it creates a mutex named "Global\\RaccoonStealer_v2_Mutex", drops a file to %APPDATA%\RaccoonStealer\config.dat, and beacons to a known C2 IP. The strongest YARA candidate string from this report is:
You are clustering a set of 30 malware samples. Fifteen samples share a C2 domain. A different set of 12 samples share a unique anti-VM evasion function. Five samples overlap between both clusters. What does this pattern suggest?
Your YARA rule contains only the condition `$string and $string2`. After scanning a clean file corpus, it produces 200 false positives. What is the most likely cause?