H3AD-REF / GUIDES / LOG PARSING & NORMALIZATION

Same Event.
Five Different Field Names.

What happens to a log before it's queryable: Syslog, CEF, LEEF, and ECS/JSON format differences, the field-mapping gotchas that break correlation across vendors, and a normalization strategy that doesn't silently lose data. This is upstream of query syntax; see KQL vs XQL vs SPL for what happens once the data is already normalized.

Format Overview

Four shapes a log commonly arrives in, from loosest to most structured.

Syslog, CEF, LEEF, ECS/JSON

Syslog (RFC 3164)
The older BSD-style format: PRI, timestamp, hostname, then free-text message
└── No strict field structure inside the message itself; everything past the header is just text

Syslog (RFC 5424)
The newer structured version: adds a version field, ISO 8601 timestamps with timezone, structured
data elements
└── Backward-compatible in spirit, not in exact parsing rules

CEF (Common Event Format, ArcSight / Micro Focus)
CEF:Version|Device Vendor|Device Product|Device Version|Signature ID|Name|Severity|Extension
└── Pipe-delimited header, then key=value pairs in the extension for everything else

LEEF (Log Event Extended Format, IBM QRadar)
LEEF:Version|Vendor|Product|Version|EventID|key=value	key=value...
└── Shares CEF's pipe-delimited header shape but with its own field names, then tab or key=value attributes

JSON / ECS (Elastic Common Schema)
Fully structured, nested JSON: source.ip, destination.port, event.action
└── Vendor-neutral field naming by design, meant to be the normalization target, not just another source format
    // most SIEM ingestion pipelines end up mapping everything toward something ECS-shaped, whether or not they call it that

Field Mapping Gotchas

The same concept, five different names, and none of them wrong.

GOTCHA

Timestamp Format And Timezone

The single most common correlation-breaking mistake
One source logs local time with no offset, another logs UTC, a third logs epoch milliseconds. Normalize to one timezone at ingest, or every correlation across sources silently skews by however many hours separate them.
GOTCHA

Source/Destination IP Naming

src_ip, srcaddr, source.ip, s_ip — same field, four spellings
Joining data across sources requires knowing every vendor's alias for the same concept. A query written against one product's field name silently returns nothing against another's, with no error to flag the mismatch.
GOTCHA

Severity Scale Differences

0-7, 0-10, or a label — pick one and convert
Syslog severity runs 0 (Emergency) to 7 (Debug). CEF severity is typically 0 to 10. Neither maps cleanly onto a vendor's own "Critical/High/Medium/Low" labels without an explicit conversion table someone has to build and maintain.
GOTCHA

Field Type Mismatches

A string that should have been a number
A port number logged as a string in one source and an integer in another breaks numeric comparisons and range queries. It fails silently, not with an error, which makes it far harder to catch during pipeline testing.

Normalization Strategy

Decisions made once, at ingest, save every query written afterward from re-solving the same problem.

STRATEGY

Normalize At Ingest, Not At Query Time

Pay the cost once
Mapping fields once during collection is far cheaper than re-deriving the mapping in every query that touches that source. Query-time normalization also means every analyst has to know the same mapping by heart.
STRATEGY

Preserve The Raw Original

Normalization is lossy by design
Keep the raw original event alongside the normalized fields. A mapping built for the common case will occasionally miss something, and the raw log is the only way to recover it after the fact.
STRATEGY

Handle Multi-Line Events Deliberately

A naive line-based parser corrupts these silently
A stack trace or a multi-line Windows event can get split into multiple log lines by a parser that assumes one line equals one event, corrupting both the original event and whatever followed it.
STRATEGY

Version Your Schema

A silent mapping change makes historical queries lie
A normalization mapping that changes over time without a version marker means the same query can produce different results depending on when the underlying data was ingested, with nothing in the data itself explaining why.

Common Pitfalls

Mistakes that show up as a wrong answer to a query, not as an error.

PITFALL

Truncated Timestamp Precision

Losing sub-second precision costs event ordering
During a fast-moving incident, sub-second sequence can be the only way to tell which of two near-simultaneous events happened first. Truncating precision during normalization throws that ordering away.
PITFALL

Silently Dropping Unmapped Fields

A rigid schema discards exactly the anomalous field that mattered
A normalization pipeline that discards anything not in the target schema throws away unexpected fields without a trace, which is often exactly the field an investigation would have needed most.
PITFALL

Case-Sensitivity Mismatches Breaking Joins

The same problem the KQL/XQL/SPL guide flags, one level upstream
The same hostname or username logged in different cases across two sources fails to correlate unless normalization forces consistent casing before the data ever reaches a query.
PITFALL

One Mapping For Every Product Version

A vendor update can break a mapping without warning
A vendor changing their log format in a product update silently breaks a mapping built against the old format, and the failure looks like missing data rather than an obvious parsing error.
ONE LEVEL DOWNSTREAM
This guide covers normalizing a log into a consistent schema before it's queryable at all. Regex For Log Parsing covers the layer above it — writing extraction patterns against fields that already exist. If you're pulling an IP or hash out of raw text at query time instead of at ingest, that guide is the one that applies.