Network Detection: NetFlow, Zeek & IDS/IPS
A SOC cannot store full packet capture of every link forever. At enterprise scale, the traffic volume makes complete payload retention prohibitively expensive within days, sometimes hours. What defenders actually build is a layered set of telemetry sources, each trading detail for cost in a different way. This chapter walks through that stack: lightweight flow metadata, Zeek's structured protocol logs, and the signature-based and behavioral engines that turn any of it into an alert. By the end you should know which source to reach for depending on what question you are trying to answer.
Flow Data vs Full Packet Capture
NetFlow, and its more modern open standard IPFIX, describe traffic as a series of flow records rather than a stream of packets. A flow record captures metadata about a conversation between two endpoints: source and destination IP, source and destination port, protocol, start and end time, and counts of packets and bytes transferred in each direction. What it does not capture is the payload.
You know that 10.1.4.22 talked to 185.220.101.4 on TCP/443 for six seconds and exchanged 42KB, but you have no visibility into what was actually inside those packets.
- Flow record: a metadata summary of one conversation between two endpoints, no payload included.
- NetFlow: Cisco's original flow-export format, still the de facto standard for flow metadata.
- IPFIX: the modern, vendor-neutral standard that succeeded NetFlow.
- PCAP: full packet capture, every byte of every packet including payload.
Why Flow Data Is Cheap
That absence of payload is the entire point. A flow record for a conversation is a few hundred bytes regardless of how much data actually moved, because it is a summary rather than a copy.
A network pushing tens of gigabits per second generates flow data that a modest server can ingest and retain for months. The same link captured at full fidelity, header and payload both, generates a volume that most organizations can only afford to keep for a matter of days before storage costs force rotation.
What Full Packet Capture Adds
Full packet capture (PCAP) is the opposite end of the spectrum. It stores every byte of every packet that crosses the tap point: headers, payload, TLS handshake bytes, everything.
This is the only source that lets you reconstruct exactly what was sent, replay a session, extract a transferred file, or examine unencrypted payload byte for byte. It is also the most expensive telemetry an organization can collect, both in storage and in the compute needed to write and query it at line rate.
The Retention Tradeoff
The practical tradeoff shows up as a retention curve. Flow data is cheap enough to keep for six to twelve months and gives you long-horizon visibility into volumetric patterns: which hosts talk to which, how much, how often.
Full capture is precise enough to answer exactly what happened in a specific incident but expensive enough that most organizations only retain it for a rolling window of one to two weeks, sometimes triggered selectively rather than run continuously across every segment. Neither source replaces the other: flow data tells you where to look, full capture tells you what happened there.
| Attribute | Flow Data (NetFlow/IPFIX) | Full Packet Capture |
|---|---|---|
| Contents | IPs, ports, protocol, byte/packet counts, duration | Complete headers and payload |
| Record size per session | Hundreds of bytes | Full size of the traffic itself |
| Typical retention | Months | Days to a couple of weeks |
| Best for | Volumetric anomalies, beaconing, long-horizon trends | Forensic reconstruction of a specific event |
| Cost driver | Ingest and storage of small records at scale | Storage and I/O for raw traffic volume |
What Zeek Actually Logs
Zeek, formerly known as Bro, is an open-source network security monitor that sits between flow data and full packet capture in both detail and cost. Rather than storing raw packets or reducing traffic to bare metadata, Zeek parses traffic protocol by protocol and writes structured, human-readable logs describing what actually happened at the application layer.
It does not usually retain payload by default, but it extracts the meaningful fields out of that payload before discarding the raw bytes, which is what makes it so useful for detection work.
The Five Core Logs
Every other Zeek log links back to a connection via a shared UID, so conn.log is usually the starting point for an investigation before pivoting into protocol-specific detail.
| Log | What It Records | Best For |
|---|---|---|
| conn.log | Every TCP, UDP, and ICMP connection: source/destination, ports, duration, bytes transferred, connection state | Foundation log, closest thing Zeek has to a flow record, starting point before pivoting elsewhere |
| dns.log | Every DNS query and response: query name, query type, response codes, answers, TTLs | Hunting DGA activity, unusual TXT records, a beacon domain that just started resolving |
| http.log | Method, URI, user-agent, response code, content type, file sizes for request and response | HTTP-based investigation without needing the raw payload |
| ssl.log | Negotiated TLS version and cipher, the SNI value requested, certificate issuer and validity period | Inspecting the handshake even when the session's application data is encrypted |
| files.log | MIME type, size, and computed hashes for files transferred across any supported protocol | Pivoting from a suspicious binary hash back to the connection and host that pulled it down |
Why Structured Logs Beat Raw Packets Here
The reason this structured, protocol-aware approach often beats raw packets for detection work is simple: an analyst can run a query against dns.log for anomalous query patterns across a month of traffic in seconds.
Decoding and reassembling raw PCAP to answer the same question over the same time window would cost far more, in both storage and compute, than anyone wants to pay.
Signature-Based IDS
How Signature Matching Works
Signature-based intrusion detection tools, Snort and Suricata being the two most widely deployed, work by comparing observed traffic against a database of rules describing known-bad patterns. A rule might match a specific byte sequence in a packet payload associated with a particular exploit, a known malicious file hash, a C2 beacon's distinctive header structure, or a regex pattern matching a known malware family's command string.
When traffic matches a rule, the engine fires an alert, and in inline deployments can drop the packet outright.
Strengths
The strength of this approach is precision and speed. When a signature matches, you know with high confidence exactly what was detected, because the rule was written against a specific, documented threat.
Rule sets such as the Emerging Threats or Snort community rules are updated constantly as new exploits and malware families are published, so a well-maintained signature engine keeps pace with known threats relatively well. Alerts are also cheap to triage, since the rule name itself tells the analyst what matched.
Limitations
The blind spot is equally fundamental: signature-based detection can only catch what it has a signature for. A novel exploit, a slightly repacked piece of malware, or an attacker who deliberately alters a byte sequence to avoid a known pattern will simply not trigger any rule.
This is not a tuning problem that better rule-writing solves, it is a structural limitation of matching against known patterns. Attackers who know a target runs Suricata with public rule sets can and do test their tooling against those exact rule sets before deployment specifically to confirm it slips through clean.
This limitation is why signature engines are best understood as one layer rather than a complete detection strategy. They are extremely good at cheaply and reliably catching the large volume of commodity, previously-seen threats that make up most day-to-day traffic, which frees analyst attention for the harder cases that require a different detection approach entirely.
Behavioral and Anomaly-Based Detection
How Behavioral Detection Works
Behavioral, or anomaly-based, detection takes the opposite starting point from signatures. Instead of asking "does this match something known-bad," it first establishes a baseline of what normal traffic looks like for a given host, segment, or user, and then flags activity that deviates meaningfully from that baseline.
A baseline might describe which ports a server typically listens on, what volume of outbound traffic is normal for a workstation, what hours a service account typically authenticates during, or what set of external hosts a given segment usually talks to.
Once that baseline exists, the engine watches for deviation: a workstation that suddenly starts making large outbound transfers at 3am, a server that starts listening on a port it has never used before, a host that begins beaconing to an external IP at suspiciously regular intervals.
None of these observations require a known signature. The system is not asking whether this traffic matches a documented threat, it is asking whether this traffic is unusual for this environment, which is exactly the kind of question that catches genuinely novel activity a signature engine would sail past.
The Tradeoff: False Positives
The tradeoff is false positive rate. Real environments are messy. A legitimate software update, an unusual but authorized backup job, or a new business process can all produce traffic that looks anomalous against a baseline built from a quieter period.
Tuning a behavioral system well takes real time investment: baselines have to be built over a long enough window to capture normal variation, and thresholds have to be adjusted per environment rather than taken as vendor defaults. Poorly tuned behavioral detection drowns analysts in noise fast enough that they start ignoring the tool entirely.
Why Programs Run Both
Because signatures and behavioral analysis fail in opposite directions, one missing novel threats, the other generating noise on legitimate-but-unusual activity, mature detection programs run both together rather than choosing one.
Signatures handle the high-confidence, well-understood threat volume cheaply. Behavioral detection catches the threats that were designed specifically to avoid triggering a signature, at the cost of requiring ongoing tuning and a higher analyst triage burden on the alerts it does produce.
IDS vs IPS: Passive Detection vs Inline Blocking
IDS and IPS often run the exact same detection engine and rule set. What differs is where the sensor sits in the network path and what it is allowed to do with a match.
Where Each Sits
An IDS is deployed passively, watching a copy of traffic mirrored to it via a SPAN port or a network tap. It sees everything that crosses the monitored link, but because it is only looking at a copy, it can alert on what it finds but cannot stop the original traffic from reaching its destination. If an IDS detects an exploit mid-flight, the exploit has already been delivered by the time the alert fires.
An IPS sits inline, directly in the path the traffic must physically traverse to reach its destination. Because every packet passes through the IPS on its way through the network, it can make a real-time decision to drop or reject a packet that matches a rule before it ever reaches the target host. This is the meaningful capability upgrade IPS offers over IDS: the ability to actually stop an attack rather than just observe and alert on it after the fact.
| Attribute | IDS | IPS |
|---|---|---|
| Position | Passive, watching a mirrored copy via SPAN port or tap | Inline, directly in the traffic's physical path |
| Action on match | Alert only, cannot stop the original traffic | Can drop or reject the packet before it reaches the target |
| Cost of a false positive | A wasted analyst triage cycle | A potential outage, dropped legitimate traffic |
| Typical rollout | Default staging ground for new rules | Promoted here only once a rule's hit rate is confirmed clean |
Why a False Positive Means More in IPS Mode
That capability comes with real operational risk. Because an inline IPS can block traffic, a false positive is no longer just a wasted analyst triage cycle, it is an outage.
A rule that misfires on legitimate business traffic can drop a critical application connection, break a partner integration, or take down a service, and it will do so silently from the perspective of anyone not watching the IPS logs. The blast radius of a bad rule in IDS mode is analyst annoyance. The blast radius of the same bad rule in IPS mode is a production incident.
Staged Rollout
Because of that asymmetry, organizations typically stage new detection content through IDS mode first. A new rule, or a newly deployed sensor, runs passively for a defined period, sometimes weeks, while the team observes what it fires on and confirms the hit rate against real traffic is clean before promoting it to blocking mode.
Even mature IPS deployments commonly keep a subset of rules, particularly ones with any history of false positives or ones covering less-understood traffic patterns, permanently in alert-only mode rather than blocking, accepting slower response on those specific rules in exchange for not risking an outage.
Choosing the Right Telemetry Source
None of these sources compete with each other, they answer different questions at different cost points, and a mature detection program uses all of them deliberately rather than defaulting to whichever is easiest to deploy. The practical skill is matching the question you are asking to the source that can actually answer it without overspending on detail you do not need.
| Source | Reach For It When | Example Question |
|---|---|---|
| NetFlow | The question is about volume, timing, or relationship patterns, not content | Is this host beaconing on a regular interval? Which internal hosts talk to which external ranges? |
| Zeek logs | The investigation needs protocol-specific facts rather than just volume | What domain did this host resolve? What certificate did this TLS session present? |
| Full packet capture | The investigation has narrowed to a specific incident and needs byte-for-byte reconstruction | What file was dropped? What data actually left during exfiltration? |
NetFlow: Volume and Pattern Questions
NetFlow is enough when the question is fundamentally about volume, timing, or relationship patterns rather than content. Detecting a host beaconing to an external IP at suspiciously regular intervals does not require knowing what was in the payload, it requires knowing that connections happened on a repeating cadence, which is exactly what a flow record captures.
The same is true for spotting a sudden spike in outbound data volume from a segment that normally sees little traffic, or mapping which internal hosts are talking to which external ranges across a long time window. Flow data's low cost makes it the right default for this class of long-horizon, pattern-level question.
Zeek: Protocol-Specific Facts
Zeek logs are the right level of detail once the investigation needs protocol-specific facts rather than just volume. If you need to know what domain a host actually resolved, what user-agent a suspicious HTTP request carried, or what certificate a TLS session presented, that information lives in Zeek's structured logs and does not exist in a flow record at all.
This is the tier most tactical investigation and hunt work actually happens at: specific enough to answer real questions, cheap enough to query across weeks of history.
Full Packet Capture: Byte-for-Byte Reconstruction
Full packet capture earns its storage cost only when the investigation has narrowed to a specific incident and the team needs to reconstruct exactly what happened, byte for byte: extracting a dropped file for malware analysis, confirming exactly what data left in an exfiltration event, or replaying a session to settle a dispute about what a host actually sent.
Because of the storage cost involved, PCAP is generally not run as a blanket, always-on capture across every segment. It is triggered selectively, often by an alert from one of the cheaper sources, and retained only for the window an active investigation needs. The next chapter walks through exactly that kind of deep forensic reconstruction using PCAP once an incident has narrowed to a specific host and time window.
Key Takeaways
- NetFlow/IPFIX records metadata only, source/destination IP and port, byte and packet counts, duration, with no payload, which is what makes it cheap enough to retain for months.
- Full packet capture stores everything including payload, giving complete forensic fidelity at a storage cost that usually limits retention to days or a couple of weeks.
- Zeek parses traffic protocol by protocol into structured logs like conn.log, dns.log, http.log, ssl.log, and files.log, giving detail close to full capture at a fraction of the cost.
- Signature-based IDS/IPS tools like Snort and Suricata match against known-bad patterns fast and precisely, but cannot detect anything novel or deliberately altered to evade a rule.
- Behavioral detection baselines normal traffic and flags deviation, catching novel threats a signature engine misses at the cost of a higher false positive rate and ongoing tuning.
- IDS watches a mirrored copy of traffic and can only alert; IPS sits inline and can block, which is powerful but turns a false positive into a potential outage rather than just noise.
Knowledge Check
Click an answer to reveal the explanation.
An analyst wants to detect a host beaconing to an external IP at suspiciously regular intervals over the past three months. Which telemetry source is the most appropriate starting point?
A team just deployed a new custom Suricata rule intended to catch a specific C2 framework's traffic pattern. What is the recommended way to roll it out?
During an investigation, an analyst confirms via Zeek's dns.log that a host resolved a suspicious domain, then needs to determine exactly what data left the host afterward, byte for byte, to confirm exfiltration. What should they turn to next?