CHAPTER 10 40 MIN READ ADVANCED

Identity and Cloud-Native Hunting

Every technique in the first nine chapters assumes a host with a process tree, a file system, and Sysmon telemetry. Identity and cloud attacks break that assumption. There is no process to inject into when the entire attack is a stolen bearer token replayed from an attacker-controlled laptop, and no file system to scan when a role-assumption chain grants an adversary access to production data three cloud accounts away from the account they compromised. This chapter builds a hunting model for a world where the perimeter is identity and the artifact is a log line, not a process.

Entra ID AWS IAM OAuth

Why Identity Is the New Perimeter

Why the Old Model Breaks

A traditional intrusion has a shape a network defender recognizes: a dropper touches disk, a process injects into another process, a beacon calls out over a socket. An identity-centric intrusion has none of that shape. The adversary authenticates. That is the entire "exploit." CrowdStrike's annual threat reporting has tracked this for several years running: a majority of detected intrusions now involve valid credentials rather than malware, because credential theft is cheaper, faster, and far less likely to trip an EDR signature than dropping a binary.

Three Structural Facts

These make identity and cloud environments require a different hunting model than the endpoint-centric approach used in Chapters 1 through 9:

  • There is no fixed network boundary. A SaaS application, an IaaS console, and an on-prem workstation are all reachable from the same stolen session token, from any IP address on the internet. "Inside" and "outside" the network stopped being a meaningful distinction the moment authentication became the control point.
  • The artifact is a control-plane log line, not a file or a process. An adversary who assumes an IAM role, grants an OAuth application access to a mailbox, or adds themselves to a privileged group leaves a JSON audit record, not a PE file on disk. Hunting here means reading structured API call logs, not parsing Sysmon.
  • Blast radius is defined by trust relationships, not network segments. A single compromised service principal, break-glass account, or federated identity provider can cascade into every tenant, subscription, or workspace that trusts it. VLANs and firewalls do not describe that topology; role assignments and trust policies do.

Applying the Same Discipline

This chapter does not replace the ATT&CK-driven hypothesis process from Chapter 2 or the PEAK/TaHiTI lifecycle from Chapter 3. It applies the same discipline to a different set of data sources: identity provider sign-in logs, cloud control-plane audit trails, and SaaS/OAuth grant records. The relevant ATT&CK techniques are T1078.004 (Valid Accounts: Cloud Accounts), T1021.007/.008 (Remote Services: Cloud Services / Direct Cloud VM Connections), and T1550.001 (Use Alternate Authentication Material: Application Access Token). TTPHUNT carries full hunting playbooks for the first two; this chapter builds the identity-specific hunting model that sits underneath them.

Note: This chapter uses Microsoft Entra ID (formerly Azure AD) and AWS IAM as worked examples because they cover the two most common enterprise identity/cloud stacks. The hunting model, not the specific product, is the transferable skill: every major identity provider and cloud platform exposes an equivalent sign-in log, risk-scoring engine, and control-plane audit trail.

Hunting Signals in Microsoft Entra ID

Risk Detections as Hunt Leads

Entra ID Protection generates risk detections against every sign-in and every user, based on signals Microsoft correlates across its global identity telemetry. These detections are hunt leads, not verdicts: a risk score is a hypothesis generator, and the hunter's job is to investigate whether the underlying sign-in was actually the adversary or a benign edge case the algorithm has not learned yet.

Risk Detection What It Means Common False Positive Source
Anonymous IP address Sign-in from a Tor exit node or a known anonymizing VPN service Privacy-conscious users on personal devices; rarely legitimate on corporate accounts
Atypical travel Sign-in from a location statistically inconsistent with the user's recent history, without necessarily being physically impossible New office locations, first use of a corporate VPN gateway, contractors working from a new country
Impossible travel Two sign-ins from locations that could not both be reached by the same person given the time elapsed between them Corporate VPN egress points that do not match the user's physical location; mobile carrier IP relocation
Malicious IP address Sign-in from an IP address with a known association to malicious activity, per Microsoft's threat intelligence Shared NAT/CGNAT IP ranges where a malicious actor and a legitimate user share an egress IP
Unfamiliar sign-in properties A sign-in that deviates from the user's established pattern of device, browser, OS, or ISP New device purchase, browser reinstall, first sign-in after a long absence (parental/medical leave)
Token issuer anomaly A SAML token was issued by a token issuer that does not match the expected federation configuration, a strong indicator of a forged token (Golden SAML) Very low false-positive rate; recent federation reconfiguration is the only common benign cause

The Sign-In Log as Hunting Surface

The sign-in log itself (Entra ID's SigninLogs table in Log Analytics/Sentinel) is the primary hunting surface. Fields worth building hypotheses around: ResultType (the exact auth error/success code), ConditionalAccessStatus, AppDisplayName, ClientAppUsed (legacy authentication protocols like IMAP and POP bypass modern MFA and are a high-value hunt target on their own), DeviceDetail.trustType, and RiskLevelDuringSignIn.

KQL
// Hunt: legacy authentication protocol use bypassing modern MFA/Conditional Access
SigninLogs
| where ClientAppUsed in ("Other clients", "IMAP4", "POP3", "Authenticated SMTP", "Exchange ActiveSync")
| where ResultType == 0
| summarize SignInCount = count(), Apps = make_set(AppDisplayName), IPs = make_set(IPAddress)
    by UserPrincipalName
| where SignInCount > 0
| order by SignInCount desc
Warning: Legacy authentication protocols cannot enforce Conditional Access or per-app MFA in the same way modern auth does. An adversary in possession of a valid username/password pair but no second factor will frequently probe legacy protocols first, specifically because they are the path of least resistance around an otherwise well-configured tenant. Any legacy-protocol sign-in success on an account that has MFA enabled is a hunt lead worth investigating, not just an inventory item to note for deprecation.

Hunting Signals in AWS IAM and CloudTrail

Why AWS Needs a Different Approach

AWS has no single "sign-in log" equivalent to Entra ID's SigninLogs. Identity activity is scattered across CloudTrail management events, and the hunter has to know which event names matter. GuardDuty pre-correlates a subset of these into managed findings, which are excellent hunt leads but do not replace direct CloudTrail hunting: GuardDuty only detects patterns its managed detectors already know about.

Key Signals to Hunt

Signal Source What To Hunt For
ConsoleLogin CloudTrail Root account logins (should be zero in a healthy account), console logins without MFA, logins immediately followed by IAM policy changes
AssumeRole / AssumeRoleWithSAML / AssumeRoleWithWebIdentity CloudTrail Long or unusual role-assumption chains (role A assumes role B assumes role C), cross-account assumption to an account outside the organization, assumption from an unfamiliar source IP
CreateAccessKey / GetSessionToken CloudTrail New long-lived access keys created for an existing IAM user shortly after other suspicious activity; session tokens requested from an IP inconsistent with the user's normal pattern
UnauthorizedAccess:IAMUser/InstanceCredentialExfiltration GuardDuty EC2 instance metadata service (IMDS) credentials used from outside the instance, a strong signal of SSRF-driven credential theft (T1552.005)
UnauthorizedAccess:IAMUser/TorIPCaller / Recon:IAMUser/TorIPCaller GuardDuty API calls made using IAM credentials from a Tor exit node
IAMUser/ConsoleLoginSuccess.B GuardDuty Console login pattern anomalous relative to the account's history, GuardDuty's closest analog to Entra's impossible-travel detection

Tracing Role-Assumption Chains

The single highest-value hunt in AWS identity is tracing role-assumption chains. A compromised low-privilege IAM user is only dangerous in proportion to what it can assume its way into.

1
Role A
Compromised low-privilege IAM user
→
2
Role B
Assumed via a chained AssumeRole call
→
3
Role C
Terminal role, the actual blast radius

A hunt hypothesis worth running quarterly regardless of any specific trigger: enumerate every AssumeRole event in the last 90 days, build a graph of which principals assumed which roles, and flag any chain longer than two hops or any chain that crosses an account boundary the security team was not aware of.

Tip: CloudTrail records the sourceIPAddress and userAgent for every API call, including role assumptions made entirely from the AWS SDK with no browser involved. An AssumeRole call with a userAgent of a generic scripting library (Python botocore, raw curl) from an IP outside your known CI/CD or VPN ranges is a stronger lead than the GuardDuty finding alone, because it is not filtered through a managed detector's threshold.

OAuth and SaaS Token Abuse

OAuth as a Persistence Mechanism

Once an organization's identity provider federates into dozens of SaaS applications, the OAuth consent grant becomes a persistence mechanism that survives a password reset. An adversary who tricks a user into approving a malicious OAuth application (an "illicit consent grant" attack) obtains a refresh token scoped to whatever permissions the app requested (mail read, file access, calendar access) that keeps working even after the victim changes their password, because OAuth tokens are not password-derived.

This maps to T1550.001, Use Alternate Authentication Material: Application Access Token. The hunting surface is the tenant's OAuth application/service principal inventory and consent audit log, not the sign-in log.

Three Hunts to Run

  • Hunt newly registered applications with high-privilege scopes. An app requesting Mail.Read, Files.ReadWrite.All, or offline_access (which grants a refresh token) within days of being registered, especially one published by an unverified developer, is a lead worth manual review regardless of whether any user has consented yet.
  • Hunt for admin consent granted outside a change window. Tenant-wide admin consent bypasses the per-user consent prompt entirely. Any admin-consent event that does not correlate to a documented procurement or IT change ticket is worth investigating immediately, not queuing for the next audit cycle.
  • Hunt for consent grants immediately following a phishing-adjacent sign-in. The temporal correlation between a suspicious sign-in risk event and a consent grant within the same session is a strong compound signal.

A common attack chain behind that third hunt:

1
Phishing Link
Email delivers a link to a legitimate Microsoft/Google consent page for an attacker-registered app
→
2
User Approves
Grants the requested scopes to the malicious app
→
3
Persistent Access
Attacker never needs the user's password again

Worked Example

Worked Example:

A hunter reviewing the OAuth application inventory in a mid-size tenant finds an app named "Document Viewer" registered six days earlier, published by a developer with no verified domain, requesting Mail.Read and Files.Read.All with offline_access. Three users have granted consent. Cross-referencing sign-in logs, all three consent events occurred within two minutes of a sign-in flagged with the "unfamiliar sign-in properties" risk detection. The pattern is consistent with a phishing campaign that delivered a consent-phishing link rather than a credential-harvesting page, explaining why no risky password-reset or MFA-fatigue activity was ever observed on these accounts. Revoking the app's consent grants and disabling the service principal cuts off access immediately, without requiring a password reset for any of the three users.

Direct Cloud VM Connections and Hybrid Blind Spots

A Genuine Blind Spot

T1021.008, Direct Cloud VM Connections, is a lateral movement technique unique to IaaS: an adversary with cloud control-plane credentials can use native connectivity services (AWS Systems Manager Session Manager, Azure Bastion, Google Cloud IAP tunneling) to open an interactive session directly to a VM without ever touching the corporate network or triggering a traditional RDP/SSH detection.

This is a genuine hunting blind spot in organizations that built their detection stack around network-based lateral movement (RDP event 4624/4625 Logon Type 10, SSH auth logs). A session opened through AWS SSM Session Manager or Azure Bastion generates a control-plane API event (StartSession in CloudTrail, or an Azure Activity Log entry for the Bastion resource), not a network-visible RDP/SSH handshake. A network-centric hunt program that never extends into the cloud control plane will never see this technique at all.

Where the Hunt Signal Lives, by Platform

Platform Service Where the Hunt Signal Lives
AWS Systems Manager Session Manager CloudTrail StartSession / TerminateSession events; SSM session logs in S3/CloudWatch if session logging is enabled
Azure Azure Bastion / Just-In-Time VM Access Azure Activity Log entries for the Bastion resource; JIT access request approvals in Microsoft Defender for Cloud
GCP Identity-Aware Proxy (IAP) TCP forwarding Cloud Audit Logs for the IAP tunnel resource; VPC Flow Logs showing the IAP relay IP range as the source
Note: Session logging for these services is frequently disabled by default or configured as opt-in. Before writing a single hunt query against SSM or Bastion session data, verify the logging is actually enabled and retained: the SANS PAM "Equip" stage from Chapter 3 applies directly here. A hunt that returns zero results because logging was never turned on is a false negative, not a clean bill of health.

Building Cloud and Identity Hunt Hypotheses

Applying ABLE to Cloud

The ABLE framework from Chapter 2 (Actor, Behavior, Location, Evidence) applies to identity and cloud hunts exactly as it does to endpoint hunts. The difference is that "Location" now means a tenant, subscription, or SaaS application rather than a subnet, and "Evidence" is a control-plane log field rather than a Sysmon event ID.

Hunt Plan Template

TEMPLATE
HUNT PLAN
Analyst: [name]
Hypothesis: An adversary who compromised a low-privilege IAM user credential
  used AssumeRole chaining to escalate into a role with S3 read access on a
  data-classification-restricted bucket (T1078.004)
ATT&CK: T1078.004 - Valid Accounts: Cloud Accounts

Scope:
- Accounts: Production AWS Organization, all member accounts
- Time window: Last 90 days
- Data sources: CloudTrail management events (AssumeRole, AssumeRoleWithSAML),
  S3 data-event logging on the restricted bucket, GuardDuty findings

Expected evidence if hypothesis TRUE:
- AssumeRole chain of 2+ hops terminating in a role with s3:GetObject
  on the restricted bucket
- Source IP or user agent on the final AssumeRole call inconsistent with
  the known CI/CD or human-user baseline for that role
- S3 data-event log showing GetObject calls shortly after the role assumption

Expected evidence if hypothesis FALSE:
- All AssumeRole chains into the restricted role originate from documented
  CI/CD service accounts with consistent source IPs and user agents

Planned queries:
1. Extract all AssumeRole/AssumeRoleWithSAML events, 90-day window
2. Build a directed graph: source principal -> assumed role
3. Filter to chains of length >= 2 terminating in roles with access to the
   restricted bucket's resource policy
4. Cross-reference terminal AssumeRole source IP/user agent against baseline

Three Ready-to-Adapt Hypotheses

For a first pass at an identity/cloud hunt program:

  • Legacy auth bypass: "An adversary with a valid password but no MFA device used a legacy authentication protocol (IMAP/POP/ActiveSync) to bypass Conditional Access." (T1078.004). Test against SigninLogs filtered on ClientAppUsed.
  • Consent phishing: "An adversary registered or leveraged a malicious OAuth application to obtain a persistent refresh token without needing the victim's password." (T1550.001). Test against the tenant's app consent audit log.
  • Cloud lateral movement blind spot: "An adversary used SSM Session Manager or Azure Bastion to move laterally to a production VM without generating a network-visible RDP/SSH event." (T1021.008). Test against CloudTrail StartSession events correlated with the absence of any corresponding network-layer logon.

Common False Positives and Pitfalls

Pitfalls Unique to Identity Hunting

Identity and cloud hunting has a different false-positive profile than endpoint hunting, and analysts who transfer endpoint intuition directly will burn credibility with noisy escalations.

Pitfall Why It Happens How To Handle It
Corporate VPN egress triggers impossible/atypical travel The identity provider sees the VPN's public egress IP location, not the user's actual location, and the two may be different countries Register the VPN's egress IP ranges as named/trusted locations in the identity provider before treating this detection as high-confidence
Service principals and CI/CD accounts flagged as anomalous Automated principals do not have a stable "typical behavior" baseline the way a human user does across weeks of activity, so behavioral risk scoring performs worse on them Maintain a separate baseline and hunt cadence for non-human identities; do not apply the same risk-score threshold used for human users
Break-glass / emergency-access accounts trigger every detection in the book These accounts are rarely used, by design, so any usage is statistically anomalous even when it is the legitimate emergency procedure being executed Alert on break-glass account usage unconditionally (any use is worth a phone call to confirm), rather than relying on risk scoring to distinguish legitimate from malicious use
Shared NAT/CGNAT ranges pollute IP-based reputation signals Large mobile carriers and some corporate networks share a small pool of public IPs across thousands of unrelated users, so a malicious IP reputation hit may reflect a different user entirely Weight IP reputation lower than behavioral/device-based signals when the source IP falls in a known CGNAT or mobile carrier range
Warning: Do not build an identity hunt program that only reacts to vendor risk scores. Entra ID Protection and GuardDuty are excellent triage layers, but both only detect patterns their vendor has already modeled. AssumeRole chain graphing, legacy-auth hunting, and OAuth consent review in this chapter all surface activity that never generates a managed finding at all, because no vendor detector is looking for that specific pattern in your specific environment.

Key Takeaways

  • Identity-centric intrusions have no process tree or file system artifact. The hunting surface is control-plane log lines: sign-in logs, CloudTrail management events, and OAuth consent audit records.
  • Entra ID Protection risk detections (impossible travel, atypical travel, anonymous IP, token issuer anomaly) are hunt leads to investigate, not verdicts to act on directly: each has a known false-positive source worth ruling out first.
  • AWS identity hunting centers on tracing AssumeRole chains. A compromised low-privilege IAM user is dangerous only in proportion to what roles it can assume its way into.
  • OAuth consent-grant abuse (T1550.001) creates persistence that survives a password reset, because refresh tokens are not password-derived. Hunt the app inventory and consent log, not just the sign-in log.
  • Direct Cloud VM Connections (T1021.008) via SSM Session Manager, Azure Bastion, or GCP IAP is a genuine blind spot for network-centric lateral movement detection, since these sessions never generate a traditional RDP/SSH event.
  • Non-human identities (service principals, CI/CD accounts) and break-glass accounts need separate baselines and hunt cadences: applying human-user risk thresholds to them produces either alert fatigue or missed abuse.

Knowledge Check

Click an answer to reveal the explanation.

Q1. Why does a stolen bearer token or OAuth refresh token survive a victim's password reset?

  • A. Because tokens are cached locally on the identity provider's servers and cannot be revoked
  • B. Because OAuth tokens are not password-derived; they are independent credentials issued at consent time and remain valid until explicitly revoked
  • C. Because password resets only apply to the primary authentication factor, never to secondary factors
  • D. Because tokens are stored in Active Directory and replicate independently of the user object

Q2. A hunter sees repeated Entra ID "atypical travel" risk detections for employees using the corporate VPN. What is the correct response?

  • A. Disable the atypical travel detection tenant-wide since it is clearly broken
  • B. Force a password reset for every affected user immediately
  • C. Register the VPN's egress IP ranges as named/trusted locations so the detection stops false-flagging known-good egress points
  • D. Ignore all future atypical travel detections tenant-wide

Q3. Why is tracing AssumeRole chains a high-value AWS identity hunt?

  • A. Because AssumeRole calls are always malicious and require no further investigation
  • B. Because a compromised low-privilege IAM user's actual blast radius is defined by every role it can assume its way into, which a single-event view of CloudTrail will not reveal
  • C. Because GuardDuty cannot see AssumeRole events at all
  • D. Because role assumption is deprecated in favor of long-lived access keys

Q4. What makes T1021.008 (Direct Cloud VM Connections) a blind spot for network-centric detection programs?

  • A. It uses encrypted RDP, which network sensors cannot decrypt
  • B. Services like SSM Session Manager and Azure Bastion open the session through the cloud control plane, generating an API event rather than a network-visible RDP/SSH handshake
  • C. It only works on-premises, so cloud logging never captures it
  • D. It requires physical access to the data center

Q5. Why do service principals and CI/CD accounts need a separate hunting baseline from human users?

  • A. Because non-human identities are immune to compromise
  • B. Because behavioral risk scoring relies on a stable pattern of "typical" activity built over time, and automated principals do not exhibit the same kind of consistent, individually-distinguishable behavior a human user does
  • C. Because service principals cannot generate CloudTrail or sign-in log events
  • D. Because non-human identities are always excluded from Conditional Access policies