Identity and Cloud-Native Hunting
Every technique in the first nine chapters assumes a host with a process tree, a file system, and Sysmon telemetry. Identity and cloud attacks break that assumption. There is no process to inject into when the entire attack is a stolen bearer token replayed from an attacker-controlled laptop, and no file system to scan when a role-assumption chain grants an adversary access to production data three cloud accounts away from the account they compromised. This chapter builds a hunting model for a world where the perimeter is identity and the artifact is a log line, not a process.
Why Identity Is the New Perimeter
Why the Old Model Breaks
A traditional intrusion has a shape a network defender recognizes: a dropper touches disk, a process injects into another process, a beacon calls out over a socket. An identity-centric intrusion has none of that shape. The adversary authenticates. That is the entire "exploit." CrowdStrike's annual threat reporting has tracked this for several years running: a majority of detected intrusions now involve valid credentials rather than malware, because credential theft is cheaper, faster, and far less likely to trip an EDR signature than dropping a binary.
Three Structural Facts
These make identity and cloud environments require a different hunting model than the endpoint-centric approach used in Chapters 1 through 9:
- There is no fixed network boundary. A SaaS application, an IaaS console, and an on-prem workstation are all reachable from the same stolen session token, from any IP address on the internet. "Inside" and "outside" the network stopped being a meaningful distinction the moment authentication became the control point.
- The artifact is a control-plane log line, not a file or a process. An adversary who assumes an IAM role, grants an OAuth application access to a mailbox, or adds themselves to a privileged group leaves a JSON audit record, not a PE file on disk. Hunting here means reading structured API call logs, not parsing Sysmon.
- Blast radius is defined by trust relationships, not network segments. A single compromised service principal, break-glass account, or federated identity provider can cascade into every tenant, subscription, or workspace that trusts it. VLANs and firewalls do not describe that topology; role assignments and trust policies do.
Applying the Same Discipline
This chapter does not replace the ATT&CK-driven hypothesis process from Chapter 2 or the PEAK/TaHiTI lifecycle from Chapter 3. It applies the same discipline to a different set of data sources: identity provider sign-in logs, cloud control-plane audit trails, and SaaS/OAuth grant records. The relevant ATT&CK techniques are T1078.004 (Valid Accounts: Cloud Accounts), T1021.007/.008 (Remote Services: Cloud Services / Direct Cloud VM Connections), and T1550.001 (Use Alternate Authentication Material: Application Access Token). TTPHUNT carries full hunting playbooks for the first two; this chapter builds the identity-specific hunting model that sits underneath them.
Hunting Signals in Microsoft Entra ID
Risk Detections as Hunt Leads
Entra ID Protection generates risk detections against every sign-in and every user, based on signals Microsoft correlates across its global identity telemetry. These detections are hunt leads, not verdicts: a risk score is a hypothesis generator, and the hunter's job is to investigate whether the underlying sign-in was actually the adversary or a benign edge case the algorithm has not learned yet.
| Risk Detection | What It Means | Common False Positive Source |
|---|---|---|
| Anonymous IP address | Sign-in from a Tor exit node or a known anonymizing VPN service | Privacy-conscious users on personal devices; rarely legitimate on corporate accounts |
| Atypical travel | Sign-in from a location statistically inconsistent with the user's recent history, without necessarily being physically impossible | New office locations, first use of a corporate VPN gateway, contractors working from a new country |
| Impossible travel | Two sign-ins from locations that could not both be reached by the same person given the time elapsed between them | Corporate VPN egress points that do not match the user's physical location; mobile carrier IP relocation |
| Malicious IP address | Sign-in from an IP address with a known association to malicious activity, per Microsoft's threat intelligence | Shared NAT/CGNAT IP ranges where a malicious actor and a legitimate user share an egress IP |
| Unfamiliar sign-in properties | A sign-in that deviates from the user's established pattern of device, browser, OS, or ISP | New device purchase, browser reinstall, first sign-in after a long absence (parental/medical leave) |
| Token issuer anomaly | A SAML token was issued by a token issuer that does not match the expected federation configuration, a strong indicator of a forged token (Golden SAML) | Very low false-positive rate; recent federation reconfiguration is the only common benign cause |
The Sign-In Log as Hunting Surface
The sign-in log itself (Entra ID's SigninLogs table in Log Analytics/Sentinel) is the primary hunting surface. Fields worth building hypotheses around: ResultType (the exact auth error/success code), ConditionalAccessStatus, AppDisplayName, ClientAppUsed (legacy authentication protocols like IMAP and POP bypass modern MFA and are a high-value hunt target on their own), DeviceDetail.trustType, and RiskLevelDuringSignIn.
// Hunt: legacy authentication protocol use bypassing modern MFA/Conditional Access
SigninLogs
| where ClientAppUsed in ("Other clients", "IMAP4", "POP3", "Authenticated SMTP", "Exchange ActiveSync")
| where ResultType == 0
| summarize SignInCount = count(), Apps = make_set(AppDisplayName), IPs = make_set(IPAddress)
by UserPrincipalName
| where SignInCount > 0
| order by SignInCount desc
Hunting Signals in AWS IAM and CloudTrail
Why AWS Needs a Different Approach
AWS has no single "sign-in log" equivalent to Entra ID's SigninLogs. Identity activity is scattered across CloudTrail management events, and the hunter has to know which event names matter. GuardDuty pre-correlates a subset of these into managed findings, which are excellent hunt leads but do not replace direct CloudTrail hunting: GuardDuty only detects patterns its managed detectors already know about.
Key Signals to Hunt
| Signal | Source | What To Hunt For |
|---|---|---|
| ConsoleLogin | CloudTrail | Root account logins (should be zero in a healthy account), console logins without MFA, logins immediately followed by IAM policy changes |
| AssumeRole / AssumeRoleWithSAML / AssumeRoleWithWebIdentity | CloudTrail | Long or unusual role-assumption chains (role A assumes role B assumes role C), cross-account assumption to an account outside the organization, assumption from an unfamiliar source IP |
| CreateAccessKey / GetSessionToken | CloudTrail | New long-lived access keys created for an existing IAM user shortly after other suspicious activity; session tokens requested from an IP inconsistent with the user's normal pattern |
| UnauthorizedAccess:IAMUser/InstanceCredentialExfiltration | GuardDuty | EC2 instance metadata service (IMDS) credentials used from outside the instance, a strong signal of SSRF-driven credential theft (T1552.005) |
| UnauthorizedAccess:IAMUser/TorIPCaller / Recon:IAMUser/TorIPCaller | GuardDuty | API calls made using IAM credentials from a Tor exit node |
| IAMUser/ConsoleLoginSuccess.B | GuardDuty | Console login pattern anomalous relative to the account's history, GuardDuty's closest analog to Entra's impossible-travel detection |
Tracing Role-Assumption Chains
The single highest-value hunt in AWS identity is tracing role-assumption chains. A compromised low-privilege IAM user is only dangerous in proportion to what it can assume its way into.
A hunt hypothesis worth running quarterly regardless of any specific trigger: enumerate every AssumeRole event in the last 90 days, build a graph of which principals assumed which roles, and flag any chain longer than two hops or any chain that crosses an account boundary the security team was not aware of.
sourceIPAddress and userAgent for every API call, including role assumptions made entirely from the AWS SDK with no browser involved. An AssumeRole call with a userAgent of a generic scripting library (Python botocore, raw curl) from an IP outside your known CI/CD or VPN ranges is a stronger lead than the GuardDuty finding alone, because it is not filtered through a managed detector's threshold.
OAuth and SaaS Token Abuse
OAuth as a Persistence Mechanism
Once an organization's identity provider federates into dozens of SaaS applications, the OAuth consent grant becomes a persistence mechanism that survives a password reset. An adversary who tricks a user into approving a malicious OAuth application (an "illicit consent grant" attack) obtains a refresh token scoped to whatever permissions the app requested (mail read, file access, calendar access) that keeps working even after the victim changes their password, because OAuth tokens are not password-derived.
This maps to T1550.001, Use Alternate Authentication Material: Application Access Token. The hunting surface is the tenant's OAuth application/service principal inventory and consent audit log, not the sign-in log.
Three Hunts to Run
- Hunt newly registered applications with high-privilege scopes. An app requesting
Mail.Read,Files.ReadWrite.All, oroffline_access(which grants a refresh token) within days of being registered, especially one published by an unverified developer, is a lead worth manual review regardless of whether any user has consented yet. - Hunt for admin consent granted outside a change window. Tenant-wide admin consent bypasses the per-user consent prompt entirely. Any admin-consent event that does not correlate to a documented procurement or IT change ticket is worth investigating immediately, not queuing for the next audit cycle.
- Hunt for consent grants immediately following a phishing-adjacent sign-in. The temporal correlation between a suspicious sign-in risk event and a consent grant within the same session is a strong compound signal.
A common attack chain behind that third hunt:
Worked Example
A hunter reviewing the OAuth application inventory in a mid-size tenant finds an app named "Document Viewer" registered six days earlier, published by a developer with no verified domain, requesting Mail.Read and Files.Read.All with offline_access. Three users have granted consent. Cross-referencing sign-in logs, all three consent events occurred within two minutes of a sign-in flagged with the "unfamiliar sign-in properties" risk detection. The pattern is consistent with a phishing campaign that delivered a consent-phishing link rather than a credential-harvesting page, explaining why no risky password-reset or MFA-fatigue activity was ever observed on these accounts. Revoking the app's consent grants and disabling the service principal cuts off access immediately, without requiring a password reset for any of the three users.
Direct Cloud VM Connections and Hybrid Blind Spots
A Genuine Blind Spot
T1021.008, Direct Cloud VM Connections, is a lateral movement technique unique to IaaS: an adversary with cloud control-plane credentials can use native connectivity services (AWS Systems Manager Session Manager, Azure Bastion, Google Cloud IAP tunneling) to open an interactive session directly to a VM without ever touching the corporate network or triggering a traditional RDP/SSH detection.
This is a genuine hunting blind spot in organizations that built their detection stack around network-based lateral movement (RDP event 4624/4625 Logon Type 10, SSH auth logs). A session opened through AWS SSM Session Manager or Azure Bastion generates a control-plane API event (StartSession in CloudTrail, or an Azure Activity Log entry for the Bastion resource), not a network-visible RDP/SSH handshake. A network-centric hunt program that never extends into the cloud control plane will never see this technique at all.
Where the Hunt Signal Lives, by Platform
| Platform | Service | Where the Hunt Signal Lives |
|---|---|---|
| AWS | Systems Manager Session Manager | CloudTrail StartSession / TerminateSession events; SSM session logs in S3/CloudWatch if session logging is enabled |
| Azure | Azure Bastion / Just-In-Time VM Access | Azure Activity Log entries for the Bastion resource; JIT access request approvals in Microsoft Defender for Cloud |
| GCP | Identity-Aware Proxy (IAP) TCP forwarding | Cloud Audit Logs for the IAP tunnel resource; VPC Flow Logs showing the IAP relay IP range as the source |
Building Cloud and Identity Hunt Hypotheses
Applying ABLE to Cloud
The ABLE framework from Chapter 2 (Actor, Behavior, Location, Evidence) applies to identity and cloud hunts exactly as it does to endpoint hunts. The difference is that "Location" now means a tenant, subscription, or SaaS application rather than a subnet, and "Evidence" is a control-plane log field rather than a Sysmon event ID.
Hunt Plan Template
HUNT PLAN
Analyst: [name]
Hypothesis: An adversary who compromised a low-privilege IAM user credential
used AssumeRole chaining to escalate into a role with S3 read access on a
data-classification-restricted bucket (T1078.004)
ATT&CK: T1078.004 - Valid Accounts: Cloud Accounts
Scope:
- Accounts: Production AWS Organization, all member accounts
- Time window: Last 90 days
- Data sources: CloudTrail management events (AssumeRole, AssumeRoleWithSAML),
S3 data-event logging on the restricted bucket, GuardDuty findings
Expected evidence if hypothesis TRUE:
- AssumeRole chain of 2+ hops terminating in a role with s3:GetObject
on the restricted bucket
- Source IP or user agent on the final AssumeRole call inconsistent with
the known CI/CD or human-user baseline for that role
- S3 data-event log showing GetObject calls shortly after the role assumption
Expected evidence if hypothesis FALSE:
- All AssumeRole chains into the restricted role originate from documented
CI/CD service accounts with consistent source IPs and user agents
Planned queries:
1. Extract all AssumeRole/AssumeRoleWithSAML events, 90-day window
2. Build a directed graph: source principal -> assumed role
3. Filter to chains of length >= 2 terminating in roles with access to the
restricted bucket's resource policy
4. Cross-reference terminal AssumeRole source IP/user agent against baseline
Three Ready-to-Adapt Hypotheses
For a first pass at an identity/cloud hunt program:
- Legacy auth bypass: "An adversary with a valid password but no MFA device used a legacy authentication protocol (IMAP/POP/ActiveSync) to bypass Conditional Access." (T1078.004). Test against
SigninLogsfiltered onClientAppUsed. - Consent phishing: "An adversary registered or leveraged a malicious OAuth application to obtain a persistent refresh token without needing the victim's password." (T1550.001). Test against the tenant's app consent audit log.
- Cloud lateral movement blind spot: "An adversary used SSM Session Manager or Azure Bastion to move laterally to a production VM without generating a network-visible RDP/SSH event." (T1021.008). Test against CloudTrail
StartSessionevents correlated with the absence of any corresponding network-layer logon.
Common False Positives and Pitfalls
Pitfalls Unique to Identity Hunting
Identity and cloud hunting has a different false-positive profile than endpoint hunting, and analysts who transfer endpoint intuition directly will burn credibility with noisy escalations.
| Pitfall | Why It Happens | How To Handle It |
|---|---|---|
| Corporate VPN egress triggers impossible/atypical travel | The identity provider sees the VPN's public egress IP location, not the user's actual location, and the two may be different countries | Register the VPN's egress IP ranges as named/trusted locations in the identity provider before treating this detection as high-confidence |
| Service principals and CI/CD accounts flagged as anomalous | Automated principals do not have a stable "typical behavior" baseline the way a human user does across weeks of activity, so behavioral risk scoring performs worse on them | Maintain a separate baseline and hunt cadence for non-human identities; do not apply the same risk-score threshold used for human users |
| Break-glass / emergency-access accounts trigger every detection in the book | These accounts are rarely used, by design, so any usage is statistically anomalous even when it is the legitimate emergency procedure being executed | Alert on break-glass account usage unconditionally (any use is worth a phone call to confirm), rather than relying on risk scoring to distinguish legitimate from malicious use |
| Shared NAT/CGNAT ranges pollute IP-based reputation signals | Large mobile carriers and some corporate networks share a small pool of public IPs across thousands of unrelated users, so a malicious IP reputation hit may reflect a different user entirely | Weight IP reputation lower than behavioral/device-based signals when the source IP falls in a known CGNAT or mobile carrier range |
Key Takeaways
- Identity-centric intrusions have no process tree or file system artifact. The hunting surface is control-plane log lines: sign-in logs, CloudTrail management events, and OAuth consent audit records.
- Entra ID Protection risk detections (impossible travel, atypical travel, anonymous IP, token issuer anomaly) are hunt leads to investigate, not verdicts to act on directly: each has a known false-positive source worth ruling out first.
- AWS identity hunting centers on tracing AssumeRole chains. A compromised low-privilege IAM user is dangerous only in proportion to what roles it can assume its way into.
- OAuth consent-grant abuse (T1550.001) creates persistence that survives a password reset, because refresh tokens are not password-derived. Hunt the app inventory and consent log, not just the sign-in log.
- Direct Cloud VM Connections (T1021.008) via SSM Session Manager, Azure Bastion, or GCP IAP is a genuine blind spot for network-centric lateral movement detection, since these sessions never generate a traditional RDP/SSH event.
- Non-human identities (service principals, CI/CD accounts) and break-glass accounts need separate baselines and hunt cadences: applying human-user risk thresholds to them produces either alert fatigue or missed abuse.
Knowledge Check
Click an answer to reveal the explanation.
Q1. Why does a stolen bearer token or OAuth refresh token survive a victim's password reset?
Q2. A hunter sees repeated Entra ID "atypical travel" risk detections for employees using the corporate VPN. What is the correct response?
Q3. Why is tracing AssumeRole chains a high-value AWS identity hunt?
Q4. What makes T1021.008 (Direct Cloud VM Connections) a blind spot for network-centric detection programs?
Q5. Why do service principals and CI/CD accounts need a separate hunting baseline from human users?