Regex For Logs.
Precise Enough To Trust At 3am.
A static reference for writing regex that survives contact with real log data: anchors and boundaries, the extraction patterns you reach for constantly, quantifier behavior, grouping syntax, and the PCRE vs RE2 gap that decides whether a pattern even runs on your SIEM. No live regex tester here, just the syntax and the judgment calls. For where a regex pattern actually gets embedded in a detection rule, see the Sigma Rule Writing Guide.
Anchors & Boundaries
A log line is rarely just the thing you want to match. Anchors decide whether your pattern respects that or ignores it.
Start Anchor ^
^, a pattern like ERROR matches ERROR anywhere in the line, including inside a longer token such as PARSERERROR_CODE. Pin it to ^ERROR when the field you're parsing is supposed to start the line, such as a syslog severity token.End Anchor $
\d+$ grabs a number only if nothing follows it, which stops it from partially matching a longer numeric string earlier in the line.Word Boundary \b
\bcmd\b matches the standalone token cmd but not the cmd inside cmdlet or bcmd. It doesn't consume a character itself, so it costs nothing to add and it's the difference between a field match and a substring match.Why Unanchored Patterns Over-Match
^, $, a delimiter, or a \b boundary tells the engine where the field actually starts and stops instead of letting it guess.Common Extraction Patterns
The patterns that show up in almost every parser, written once so you stop re-deriving them from memory.
IPv4 Address
\b(?:\d{1,3}\.){3}\d{1,3}\b
Matches four dot-separated 1-3 digit groups. It will also match 999.999.999.999, since it checks shape, not value. If you need octets bounded to 0-255, that's a separate, longer pattern; most log parsers accept the loose version and validate the range afterward.
Email Address
\b[\w.+-]+@[\w-]+\.[a-zA-Z]{2,}\b
Covers the vast majority of addresses seen in mail headers and auth logs. A fully RFC-compliant email regex is notoriously long and rarely worth it outside of input validation on a signup form.
Hash Values By Fixed Length
\b[a-f0-9]{32}\b— MD5\b[a-f0-9]{40}\b— SHA1\b[a-f0-9]{64}\b— SHA256
Add (?i) or a case-insensitive flag if the log source ever emits uppercase hex; some AV and EDR tools do.
Domain Name
\b(?:[a-zA-Z0-9-]{1,63}\.)+[a-zA-Z]{2,}\b
Matches domains and subdomains alike. It doesn't validate against a real TLD list, so it will also match a made-up ending like .zzz; if that matters, pair it with a known-TLD check downstream.
URL
https?://[^\s"'<>]+
Simple and reliable for pulling a full URL out of a log line, since it stops at whitespace or common delimiter characters instead of trying to fully parse the path and query string grammar.
Quantifiers
How much a pattern is willing to consume before it gives up and backs off. Get this wrong and the pattern either under-matches or hangs.
Greedy: .*
key="a" key="b", a greedy pattern key="(.*)" matches from the first quote all the way to the last one, capturing a" key="b instead of just a. It grabbed the whole line because nothing told it to stop early.Lazy: .*?
key="(.*?)" stops at the first closing quote, capturing only a. Lazy quantifiers are usually what you actually want when extracting a single delimited field out of a line with repeated delimiters.Catastrophic Backtracking
(a+)+b against a string with many as and no trailing b forces the engine to try an exponential number of ways to split those as before it can declare failure. Nesting one quantifier inside another is the usual trigger; flatten the pattern or use a possessive/atomic construct where the engine supports it.Groups
Parentheses do two different jobs: grouping a sub-pattern, and capturing what it matched. Conflating them wastes cycles.
Capturing Groups ()
(\d{4})-(\d{2})-(\d{2}) captures year, month, and day as three separate groups you can reference by number afterward, such as $1, $2, $3 in a replacement or a parser's match object.Non-Capturing Groups (?:)
(?:GET|POST|PUT)\s/api groups the HTTP verb alternatives so the quantifier and the rest of the pattern apply correctly, but never allocates capture-group memory for a value you were never going to read back.Named Groups (?<name>...)
(?<src_ip>\b(?:\d{1,3}\.){3}\d{1,3}\b) reads back as match.group('src_ip') instead of match.group(1). It costs nothing extra at match time and saves the next person reading the parser from counting parentheses to figure out what group 3 was supposed to be.Engine Differences
Not every regex flavor supports the same syntax, and the pattern that works in your terminal may get rejected by the platform it's meant for.
PCRE
\1), lookahead ((?=...)), lookbehind ((?<=...)), and named groups. It trades that expressiveness for backtracking, which is exactly what makes catastrophic backtracking possible in the first place.RE2
(?=...) or \1 in PCRE has to be rewritten without them before it will run on an RE2-backed detection engine; there's no flag to enable it.grep -E vs grep -P
grep -E uses POSIX Extended Regular Expressions, no backreferences or lookaround. grep -P switches to PCRE and unlocks both, at the cost of the same backtracking exposure PCRE carries everywhere else. Know which one a script is calling before assuming a lookahead pattern will even parse.Best Practices
The habits that keep a regex parser fast, accurate, and boring in production, which is exactly what you want from it.
Anchor To Cut False Matches
^, $, or a known delimiter stops a pattern from matching a coincidental lookalike buried in an unrelated field. It's the difference between a parser that extracts the field you meant and one that occasionally extracts a substring of a free-text message.Prefer Non-Capturing Groups
(?:...) and only capturing what you'll actually read back is a real, measurable win.Test Against Edge Cases
1[.]2[.]3[.]4 or hxxp://, since threat intel feeds and analyst notes routinely defang indicators before a pattern ever sees them.Common Pitfalls
Mistakes that pass code review and pass a quick manual test, then break the first time they meet real traffic.
Unescaped Dot
. matches any character, not a literal period. 192.168.1.1 written as a pattern instead of 192\.168\.1\.1 will also match 192X168X1X1, which usually doesn't matter until it silently matches something it shouldn't have in a domain or IP field.Greedy Over-Match Across Fields
.* can swallow an entire multi-field log linemsg="(.*)" against a line with multiple quoted fields grabs from the first quote to the very last one in the line, not just the one field you meant. Swap in a lazy quantifier or a character class that excludes the delimiter, such as [^"]*, so the match stops where the field actually ends.