Query Languages: KQL, SPL, and Sigma
Every SIEM speaks its own dialect. A detection that reads cleanly in one platform's query language can look completely unfamiliar in another, even when the underlying logic is identical. Detection engineers rarely stay locked to a single tool for their whole career, and plenty work across more than one platform at the same job: a Sentinel-based SOC that also runs Splunk for a subsidiary, or a team migrating off one platform onto another mid-project. This chapter walks through the three dialects you're most likely to encounter: KQL, SPL, and Sigma, the format built specifically to move between the other two.
Why Query Language Matters
It's tempting to treat query language as a syntax preference, the way some developers argue about tabs versus spaces. In detection engineering it's closer to a practical constraint. The logic behind a detection, the condition you're trying to catch, the fields you're filtering on, the threshold that separates noise from signal, is the same no matter what platform runs it. But that logic has to be expressed in whatever language your SIEM understands, and that language shapes how you think about the problem while you're writing it.
This matters for a few concrete reasons.
- Job mobility: a detection engineer who only knows one platform's syntax is limited to jobs running that platform, and most job postings list a specific SIEM by name.
- Multi-tool environments are common, not rare: organizations that acquire other companies inherit their tooling, organizations that migrate SIEMs run two in parallel for months, and managed security service providers often support several client platforms at once.
- Community knowledge is scattered: blog posts, conference talks, and shared rule repositories are spread across all three dialects, and being able to read a KQL rule and understand what it's doing even if your shop runs Splunk saves you from reinventing detections that already exist somewhere.
KQL (Kusto Query Language)
KQL is the query language behind Microsoft Sentinel, Microsoft Defender, and the broader Azure Data Explorer family of products. If your organization runs Microsoft's security stack, KQL is the language your detections, hunting queries, and workbooks are written in.
KQL is pipe-based: you start with a table, then chain a series of operators together with the pipe character, each one narrowing, transforming, or reshaping the data before it reaches the next stage. This pipeline shape isn't unique to KQL, you'll recognize the same general pattern in SPL and in tools like PowerShell, but KQL's specific keywords and operator names are its own.
A simple illustrative example
The snippet below is a generic, illustrative example, not a real production rule, showing the general shape of a KQL query filtering a sign-in log table for failed logon events:
SigninLogs
| where ResultType != "0"
| where TimeGenerated > ago(1h)
| project TimeGenerated, UserPrincipalName, IPAddress, ResultType
| sort by TimeGenerated desc
Reading it left to right: start from the SigninLogs table, keep only rows where the result code isn't a success ("0"), keep only rows from the last hour, select down to a handful of relevant columns, and sort the output. Each where clause narrows the row set further, and project trims the columns down to what's useful to look at. That's the core rhythm of KQL: start broad, filter down, shape the output.
SPL (Splunk Search Processing Language)
SPL is Splunk's query language. It's also pipe-based, and if you've already read the KQL section above, the overall shape will feel familiar: a search starts with a source, then flows through a chain of commands separated by pipes. Where the two dialects diverge is in keyword choice, syntax conventions, and some of the underlying assumptions about how a search begins.
A Splunk search typically starts with search terms or an index specification rather than an explicit table name, and commands like where, table, and sort do work that's conceptually similar to KQL's where, project, and sort, just spelled and structured differently.
The same kind of filter, in SPL style
Here's a generic, illustrative SPL example covering the same general kind of filter as the KQL example above, failed logon events in the last hour, so you can compare the two dialects conceptually:
index=auth_logs sourcetype=signin_events
| where result_code != "0"
| where _time > relative_time(now(), "-1h")
| table _time, user, src_ip, result_code
| sort - _time
The logic is the same as the KQL version: narrow to failed results, narrow to the last hour, select the useful fields, sort by time. What differs is the vocabulary: index= and sourcetype= to scope the search, table instead of project, and SPL's particular way of expressing relative time windows. Put the two snippets side by side and the family resemblance is obvious even before you know either language well: filter, filter, shape, sort. That resemblance is exactly why being fluent in one pipe-based query language makes the next one faster to pick up.
Sigma: The Portable Format
KQL and SPL are both platform-native: a KQL query only runs against Microsoft's stack, an SPL search only runs in Splunk. That's fine if your organization only ever runs one SIEM for the rest of its existence, but detection logic is valuable independent of the platform it happens to run on, and a lot of organizations don't stay on one platform forever.
Sigma addresses that. It's a generic, YAML-based detection rule format maintained by the SigmaHQ open-source project. A Sigma rule doesn't describe KQL or SPL syntax directly, it describes the detection logic (what log source, what fields, what condition) in a structured, platform-agnostic way. Conversion tooling then translates that Sigma rule into KQL, SPL, or other backend query languages as needed. Write the logic once, generate the platform-specific query for wherever it needs to run, instead of hand-rewriting the same idea in every dialect your organization touches.
Basic structure
A Sigma rule is built from a small number of core fields:
- Title and description
logsourceblock: describes what kind of data the rule applies todetectionblock: defines the selection criteria and the condition that ties them together- Metadata: severity level, references, and similar fields
Here's a simple, illustrative Sigma rule following the same failed-logon theme as the earlier examples:
title: Multiple Failed Logon Attempts
id: 11111111-2222-3333-4444-555555555555
status: experimental
description: Detects repeated failed sign-in events for a single account within a short window
logsource:
category: authentication
product: windows
detection:
selection:
EventID: 4625
condition: selection
level: medium
The logsource block tells the conversion tooling what kind of data and product this rule expects, which helps it pick the right field mappings when generating the platform-specific query. The detection block defines one or more named selections (here, just selection, matching on the Windows failed-logon event ID) and a condition that says how those selections combine into a match. A real rule intended for production would typically add a timeframe and a count threshold to catch repeated attempts rather than a single failed logon, but the fields shown here are the structural skeleton every Sigma rule builds on.
When to Write Sigma-First vs. Native
Neither approach is universally correct. The right choice depends on what the detection is for and how long it's expected to live.
When Sigma-First Fits
Sigma-first makes sense when portability actually matters: a detection you expect to run across more than one SIEM, one you want to draw from or contribute back to community rule repositories, or one you want to version-control in a way that isn't tied to a specific platform's syntax.
When Native Fits
Native, writing directly in KQL or SPL, makes more sense when a detection needs a platform-specific function, enrichment, or data source that Sigma's common schema doesn't cleanly express, or when you're writing a quick investigative query during an incident that was never meant to become a permanent detection in the first place.
| Consideration | Sigma-first | Native (KQL / SPL) |
|---|---|---|
| Portability across SIEMs | Built in, that's the point | None, tied to one platform |
| Access to platform-specific functions | Limited to what the common schema and conversion tooling support | Full access to every native function |
| Community rule sharing | Large shared repositories to draw from and contribute to | Sharing usually means manual translation first |
| Speed for a one-off investigative query | Extra overhead for something you'll never reuse | Fastest path, write it and run it |
| Long-term maintainability across a migration | Rule logic survives a platform change | Requires a rewrite if the platform changes |
In practice, a lot of teams land somewhere in between: build the stable, well-understood detections in Sigma so they're portable and shareable, and reserve native queries for platform-specific edge cases and throwaway investigative work.
A Note on Field Normalization
Readers coming from the SOC Operations module's SIEM chapter will recognize this point, and it's worth repeating here because it applies just as much to query language as it does to SIEM architecture: none of this works without normalized field names across log sources.
The Silent Failure Mode
A detection written against a field name that doesn't exist, or that means something different, in your environment won't error out loudly. It will simply fail to match anything, quietly, regardless of whether it's written in KQL, SPL, or converted from a Sigma rule.
src_ip will convert to a query referencing whatever field your platform's mapping says corresponds to source IP. If that mapping is wrong, or that field is populated inconsistently across your log sources, the converted query runs without error and simply never fires.The query language is only as reliable as the data model underneath it, and that data model is a normalization problem, not a syntax problem.
Key Takeaways
- KQL and SPL are both pipe-based query languages, but they're platform-native: KQL runs on Microsoft Sentinel and Defender, SPL runs on Splunk, and neither runs on the other's platform.
- Despite different keywords and conventions, KQL and SPL share a common rhythm: start from a data source, filter down, shape the output, sort or aggregate as needed.
- Sigma is a YAML-based, platform-agnostic detection rule format maintained by SigmaHQ, built around
logsourceanddetectionblocks, designed to be converted into KQL, SPL, or other backend languages. - Sigma-first suits detections you want portable, shareable, or version-controlled independent of platform. Native suits detections needing platform-specific functions or quick one-off investigative queries.
- Every dialect depends on normalized, consistent field names across log sources. A mismatched or missing field causes a detection to silently fail to fire, no matter which language it's written in.
Knowledge Check
Click an answer to reveal the explanation.
What do KQL and SPL have in common structurally, even though their exact syntax differs?
What is the primary purpose of Sigma as a detection rule format?
A detection engineer writes a Sigma rule referencing a field that exists in Sigma's common schema but is populated inconsistently in the organization's actual log data. What's the most likely outcome after the rule is converted and deployed?