CHAPTER 02 30 MIN READ BEGINNER

Query Languages: KQL, SPL, and Sigma

Every SIEM speaks its own dialect. A detection that reads cleanly in one platform's query language can look completely unfamiliar in another, even when the underlying logic is identical. Detection engineers rarely stay locked to a single tool for their whole career, and plenty work across more than one platform at the same job: a Sentinel-based SOC that also runs Splunk for a subsidiary, or a team migrating off one platform onto another mid-project. This chapter walks through the three dialects you're most likely to encounter: KQL, SPL, and Sigma, the format built specifically to move between the other two.

KQLSPLSigmaquery languages
Scope note: this chapter is about reading and comparing these dialects conceptually so you can recognize them and follow along in documentation or a coworker's rule, not about becoming fluent in writing production queries in any one of them. Hands-on query construction, building filters, joins, and aggregations toward an actual detection, starts in Chapter 3.

Why Query Language Matters

It's tempting to treat query language as a syntax preference, the way some developers argue about tabs versus spaces. In detection engineering it's closer to a practical constraint. The logic behind a detection, the condition you're trying to catch, the fields you're filtering on, the threshold that separates noise from signal, is the same no matter what platform runs it. But that logic has to be expressed in whatever language your SIEM understands, and that language shapes how you think about the problem while you're writing it.

This matters for a few concrete reasons.

Why it matters:
  • Job mobility: a detection engineer who only knows one platform's syntax is limited to jobs running that platform, and most job postings list a specific SIEM by name.
  • Multi-tool environments are common, not rare: organizations that acquire other companies inherit their tooling, organizations that migrate SIEMs run two in parallel for months, and managed security service providers often support several client platforms at once.
  • Community knowledge is scattered: blog posts, conference talks, and shared rule repositories are spread across all three dialects, and being able to read a KQL rule and understand what it's doing even if your shop runs Splunk saves you from reinventing detections that already exist somewhere.
Note: knowing multiple dialects doesn't mean memorizing every function name in each one. It means recognizing the shape of the logic (filter, aggregate, join, threshold) regardless of which keywords express it, and knowing where to look up the exact syntax when you need it.

KQL (Kusto Query Language)

KQL is the query language behind Microsoft Sentinel, Microsoft Defender, and the broader Azure Data Explorer family of products. If your organization runs Microsoft's security stack, KQL is the language your detections, hunting queries, and workbooks are written in.

KQL is pipe-based: you start with a table, then chain a series of operators together with the pipe character, each one narrowing, transforming, or reshaping the data before it reaches the next stage. This pipeline shape isn't unique to KQL, you'll recognize the same general pattern in SPL and in tools like PowerShell, but KQL's specific keywords and operator names are its own.

A simple illustrative example

The snippet below is a generic, illustrative example, not a real production rule, showing the general shape of a KQL query filtering a sign-in log table for failed logon events:

SigninLogs
| where ResultType != "0"
| where TimeGenerated > ago(1h)
| project TimeGenerated, UserPrincipalName, IPAddress, ResultType
| sort by TimeGenerated desc

Reading it left to right: start from the SigninLogs table, keep only rows where the result code isn't a success ("0"), keep only rows from the last hour, select down to a handful of relevant columns, and sort the output. Each where clause narrows the row set further, and project trims the columns down to what's useful to look at. That's the core rhythm of KQL: start broad, filter down, shape the output.

SPL (Splunk Search Processing Language)

SPL is Splunk's query language. It's also pipe-based, and if you've already read the KQL section above, the overall shape will feel familiar: a search starts with a source, then flows through a chain of commands separated by pipes. Where the two dialects diverge is in keyword choice, syntax conventions, and some of the underlying assumptions about how a search begins.

A Splunk search typically starts with search terms or an index specification rather than an explicit table name, and commands like where, table, and sort do work that's conceptually similar to KQL's where, project, and sort, just spelled and structured differently.

The same kind of filter, in SPL style

Here's a generic, illustrative SPL example covering the same general kind of filter as the KQL example above, failed logon events in the last hour, so you can compare the two dialects conceptually:

index=auth_logs sourcetype=signin_events
| where result_code != "0"
| where _time > relative_time(now(), "-1h")
| table _time, user, src_ip, result_code
| sort - _time

The logic is the same as the KQL version: narrow to failed results, narrow to the last hour, select the useful fields, sort by time. What differs is the vocabulary: index= and sourcetype= to scope the search, table instead of project, and SPL's particular way of expressing relative time windows. Put the two snippets side by side and the family resemblance is obvious even before you know either language well: filter, filter, shape, sort. That resemblance is exactly why being fluent in one pipe-based query language makes the next one faster to pick up.

Sigma: The Portable Format

KQL and SPL are both platform-native: a KQL query only runs against Microsoft's stack, an SPL search only runs in Splunk. That's fine if your organization only ever runs one SIEM for the rest of its existence, but detection logic is valuable independent of the platform it happens to run on, and a lot of organizations don't stay on one platform forever.

Sigma addresses that. It's a generic, YAML-based detection rule format maintained by the SigmaHQ open-source project. A Sigma rule doesn't describe KQL or SPL syntax directly, it describes the detection logic (what log source, what fields, what condition) in a structured, platform-agnostic way. Conversion tooling then translates that Sigma rule into KQL, SPL, or other backend query languages as needed. Write the logic once, generate the platform-specific query for wherever it needs to run, instead of hand-rewriting the same idea in every dialect your organization touches.

Basic structure

A Sigma rule is built from a small number of core fields:

Core Sigma fields:
  • Title and description
  • logsource block: describes what kind of data the rule applies to
  • detection block: defines the selection criteria and the condition that ties them together
  • Metadata: severity level, references, and similar fields

Here's a simple, illustrative Sigma rule following the same failed-logon theme as the earlier examples:

title: Multiple Failed Logon Attempts
id: 11111111-2222-3333-4444-555555555555
status: experimental
description: Detects repeated failed sign-in events for a single account within a short window
logsource:
  category: authentication
  product: windows
detection:
  selection:
    EventID: 4625
  condition: selection
level: medium

The logsource block tells the conversion tooling what kind of data and product this rule expects, which helps it pick the right field mappings when generating the platform-specific query. The detection block defines one or more named selections (here, just selection, matching on the Windows failed-logon event ID) and a condition that says how those selections combine into a match. A real rule intended for production would typically add a timeframe and a count threshold to catch repeated attempts rather than a single failed logon, but the fields shown here are the structural skeleton every Sigma rule builds on.

Note: Sigma conversion depends on field mappings between Sigma's common schema and your platform's actual field names. A rule that converts cleanly assumes your environment's fields line up with what the conversion tooling expects, which connects to the normalization point covered later in this chapter.

When to Write Sigma-First vs. Native

Neither approach is universally correct. The right choice depends on what the detection is for and how long it's expected to live.

When Sigma-First Fits

Sigma-first makes sense when portability actually matters: a detection you expect to run across more than one SIEM, one you want to draw from or contribute back to community rule repositories, or one you want to version-control in a way that isn't tied to a specific platform's syntax.

When Native Fits

Native, writing directly in KQL or SPL, makes more sense when a detection needs a platform-specific function, enrichment, or data source that Sigma's common schema doesn't cleanly express, or when you're writing a quick investigative query during an incident that was never meant to become a permanent detection in the first place.

ConsiderationSigma-firstNative (KQL / SPL)
Portability across SIEMsBuilt in, that's the pointNone, tied to one platform
Access to platform-specific functionsLimited to what the common schema and conversion tooling supportFull access to every native function
Community rule sharingLarge shared repositories to draw from and contribute toSharing usually means manual translation first
Speed for a one-off investigative queryExtra overhead for something you'll never reuseFastest path, write it and run it
Long-term maintainability across a migrationRule logic survives a platform changeRequires a rewrite if the platform changes

In practice, a lot of teams land somewhere in between: build the stable, well-understood detections in Sigma so they're portable and shareable, and reserve native queries for platform-specific edge cases and throwaway investigative work.

A Note on Field Normalization

Readers coming from the SOC Operations module's SIEM chapter will recognize this point, and it's worth repeating here because it applies just as much to query language as it does to SIEM architecture: none of this works without normalized field names across log sources.

The Silent Failure Mode

A detection written against a field name that doesn't exist, or that means something different, in your environment won't error out loudly. It will simply fail to match anything, quietly, regardless of whether it's written in KQL, SPL, or converted from a Sigma rule.

Example: a Sigma rule that assumes a field called src_ip will convert to a query referencing whatever field your platform's mapping says corresponds to source IP. If that mapping is wrong, or that field is populated inconsistently across your log sources, the converted query runs without error and simply never fires.

The query language is only as reliable as the data model underneath it, and that data model is a normalization problem, not a syntax problem.

Key Takeaways

  • KQL and SPL are both pipe-based query languages, but they're platform-native: KQL runs on Microsoft Sentinel and Defender, SPL runs on Splunk, and neither runs on the other's platform.
  • Despite different keywords and conventions, KQL and SPL share a common rhythm: start from a data source, filter down, shape the output, sort or aggregate as needed.
  • Sigma is a YAML-based, platform-agnostic detection rule format maintained by SigmaHQ, built around logsource and detection blocks, designed to be converted into KQL, SPL, or other backend languages.
  • Sigma-first suits detections you want portable, shareable, or version-controlled independent of platform. Native suits detections needing platform-specific functions or quick one-off investigative queries.
  • Every dialect depends on normalized, consistent field names across log sources. A mismatched or missing field causes a detection to silently fail to fire, no matter which language it's written in.

Knowledge Check

Click an answer to reveal the explanation.

What do KQL and SPL have in common structurally, even though their exact syntax differs?

KQL and SPL each build a query as a pipeline: start from a data source, then chain operators (filters, projections, sorts, aggregations) together, typically separated by a pipe character. The exact keywords and conventions differ between the two, but the underlying pipeline shape is the same, which is part of why moving between them is easier than moving to or from an entirely different paradigm.

What is the primary purpose of Sigma as a detection rule format?

Sigma's purpose is portability: write the detection logic once in its common schema, then use conversion tooling to generate the equivalent query for whatever backend platform you need it to run on. It doesn't replace platform-native languages outright, and it doesn't automatically fix field normalization gaps, those still have to be handled through correct field mappings in your environment.

A detection engineer writes a Sigma rule referencing a field that exists in Sigma's common schema but is populated inconsistently in the organization's actual log data. What's the most likely outcome after the rule is converted and deployed?

This is the field normalization problem in practice. A converted query built on a field that's inconsistently populated, or that doesn't mean what the rule author assumed, will run without throwing any error. It simply won't match the events it was meant to catch in cases where that field's data doesn't line up, and nothing alerts you that the detection is quietly underperforming. This failure mode applies to KQL, SPL, and Sigma conversions alike: it's a data model issue, not a syntax issue.