CHAPTER 02 30 MIN READ BEGINNER

AI-Powered Threats

Generative AI did not invent phishing, fraud, or malware. What it changed is the cost of doing each of them well. A convincing pretext used to take a skilled social engineer real time to write; a cloned voice used to require studio equipment and a long audio sample; a malware variant that dodges signatures used to need a developer who understood the detection logic. This chapter walks through where AI is measurably lowering that cost for attackers, and where the capability is real versus where it is still speculation.

ai-powered threatsdeepfakesadversarial machine learningdata poisoningmitre atlas
Before you start: this chapter is about attackers using AI as a tool against your organization's normal defenses. Chapter 3 covers a different problem: attackers targeting the AI applications your organization builds or deploys. Keep that distinction in mind as you read, since the two get conflated a lot in vendor marketing.

LLM-generated phishing and social engineering

Security awareness training used to lean on a reliable set of tells: stilted grammar, generic greetings, phrasing that read like it had been translated twice. A capable LLM removes most of that signal, producing fluent, contextually appropriate prose in seconds.

Old attack patternWhat AI changes
Classic BEC (reused templates: urgent wire transfer, gift card request, fake invoice)An LLM drafts a pretext tailored to the specific sender/recipient relationship, matching tone and register per target. The template itself stops being a reliable detection surface.
Spear-phishing recon (manually researching a target's role, projects, vendors, writing style)AI tools summarize public profiles and scraped content into a usable target profile far faster than manual OSINT, making spear-phishing-grade targeting cheap enough to scale beyond high-value targets.

Neither is a new capability, BEC and OSINT-driven spear phishing already existed. What changes is the economics: effort that used to justify itself only for high-value targets now scales cheaply to many more targets.

Tells that no longer hold up reliably:
  • Awkward grammar or phrasing as a sole indicator, since fluent generated text passes this check easily
  • Generic greetings, since personalization from scraped profile data is now cheap to produce
  • Obviously templated wording, since each message can be uniquely generated rather than copy-pasted
  • Mismatched tone for the claimed sender, since a model can be prompted to imitate a specific register or prior writing sample
What still works as a detection signalWhy it survives AI-generated text
Sender authentication (SPF, DKIM, DMARC)Text quality does not change the infrastructure the message actually came from
Out-of-band verification for payment or credential requestsA phone call or a second channel confirms intent independent of how well the email reads
Behavioral and metadata anomaliesUnusual sending patterns, reply-to mismatches, and link destinations do not improve just because the prose does
Reporting culture and user skepticism toward urgency and pressureFluency does not remove the underlying social engineering pressure tactic, which is still detectable

Voice cloning and deepfakes

Voice cloning generates new audio in someone's voice from a sample of their speech. Modern tools need very little sample audio, and executives who speak on earnings calls or company videos generate that sample just by doing their jobs.

MechanismAttack pattern it enables
Voice cloning from a short public sampleAn evolution of CEO fraud/vishing: a cloned voice purporting to be an executive or vendor contact creates urgency around a wire transfer, credential reset, or policy exception. A short, urgent voicemail is often enough, it doesn't need to fool a careful listener on a long call.
Deepfake video/image generationSynthetic video of someone saying or doing something they didn't, or fabricated images supporting disinformation, fake IDs, or a social engineering chain. Quality varies; automated detection is an active research area with no fully reliable solution yet.

The psychological lever is identical to older phone-based social engineering; what changes is that "I recognized the voice" stops being a safe verification method on its own. The defensive answer is process, not better ear training: any sensitive or urgent request should go through verification that doesn't rely on the same channel or claimed identity as the request itself.

Defensive baseline: require out-of-band verification for financial or credential-sensitive requests, use a pre-agreed call-back number rather than one supplied in the request, and consider a shared verification phrase or code word for high-risk approval paths such as executive wire authorizations. None of this requires detecting that audio or video is synthetic; it just refuses to trust identity claims made over a single, unverified channel.

AI-assisted malware and offensive tooling

There's a real and growing pattern of attackers using LLMs as a coding assistant: drafting boilerplate, explaining unfamiliar APIs, converting proof-of-concept code between languages, or helping obfuscate a payload. None of this requires the model to be malicious by design, just a competent assistant that doesn't refuse the request.

ClaimEstablished?Why
AI lowers the skill floor for less sophisticated actorsYesGuardrails on mainstream LLMs push misuse toward jailbreaks and open-weight models fine-tuned to remove restrictions, marketed as "uncensored" coding assistants in criminal forums.
AI speeds up malware variant generationYesA model can rapidly produce functionally equivalent versions of the same malicious code with different structure/naming/packing, each evading a rule tuned to the previous sample. Not new in concept (polymorphism is old), just cheaper.
LLMs autonomously discover and weaponize novel zero-days without human directionNoNot a demonstrated, reproducible capability as of this writing. Treat vendor/commentary claims to the contrary skeptically.

The honest framing: AI is a capable assistant that speeds up and broadens access to offensive tooling development. It is not, currently, an autonomous threat actor.

Adversarial machine learning: evasion and data poisoning

So far this chapter has covered attackers using AI as a tool. This section covers something different: attacks that target machine learning models themselves, including the ones defenders rely on for malware classification, spam filtering, and fraud/anomaly detection.

Attack typeWhen it happensWhat it targetsHow it works
Evasion / adversarial exampleInference time, against a deployed modelCauses a specific malicious input to be misclassified as benignModifies an input (e.g. a malicious file) so it crosses the model's decision boundary without changing its actual behavior. The attacker doesn't need to know the model's internals, many techniques work by probing outputs and iterating, or transferring adversarial inputs crafted against a similar model.
Data poisoningTraining or fine-tuning timeCorrupts what the model learns, biasing future predictions or planting a backdoor triggerInfluences the data a model learns from (a public dataset, scraped content, user-submitted retraining samples) without ever breaching the model itself. Especially relevant when fine-tuning on user-supplied or less-trusted data.

This is a good point to introduce MITRE ATLAS, the Adversarial Threat Landscape for Artificial-Intelligence Systems, a MITRE-maintained knowledge base of real-world adversary tactics and techniques against AI systems, built in the same style as ATT&CK.

MITRE ATLAS at a glance:
  • Structured like ATT&CK: a matrix of tactics, with techniques and sub-techniques mapped under each, plus case studies and mitigations
  • As of its most recent 2025 update, covers around 16 tactics and roughly 80-plus techniques, including data poisoning, evasion, model extraction, and prompt-based attacks
  • Exists because ATT&CK was built for conventional IT infrastructure and doesn't natively capture attacks on models, training pipelines, or AI-specific components
  • Covered in more depth in Chapter 6, once the focus shifts to defending AI systems directly

What this means for defenders

Nothing in this chapter replaces the existing threat landscape, it sits on top of it. None of these AI-enhanced techniques change the underlying goal an attacker is trying to reach; what changes is the economics, cheaper, faster, more convincing at scale, putting attacks that used to require patient, well-resourced actors within reach of many more of them.

The practical response is process, not detection-by-eye:
  • Verification procedures that assume any request could be fabricated
  • Detection logic built on behavior and infrastructure, not surface text/audio/video polish
  • Awareness training that teaches people to recognize pressure tactics, not typos

Chapter 3 turns to a related but distinct problem: not the attacker using AI as a tool against your defenses, but the target being your own organization's AI or LLM application, through prompt injection and the broader risks of deploying LLM-powered features. Keep the two mentally separate going forward.

Key Takeaways

  • Generative AI removes the classic phishing tells, fluency, personalization, and correct grammar no longer indicate a legitimate sender, so detection has to rely on authentication, infrastructure, and behavior instead of text quality.
  • AI-assisted OSINT makes spear-phishing-grade targeting cheap enough to apply broadly, not just to high-value targets.
  • Voice cloning and deepfakes undermine identity verification based on recognizing a voice or a face; the fix is process-based out-of-band verification, not better ear training.
  • AI-assisted malware development mostly lowers the skill floor and speeds up variant generation; claims of autonomous zero-day discovery by LLMs are not an established capability.
  • Evasion attacks manipulate inputs at inference time to fool a deployed model; data poisoning corrupts what a model learns during training. They are different attacks against different points in the ML lifecycle.
  • MITRE ATLAS is a knowledge base of adversary tactics and techniques against AI systems, modeled on ATT&CK, covering roughly 16 tactics and 80-plus techniques as of its 2025 update. Chapter 6 goes deeper on defending against it.
  • AI changes the economics of existing attacks rather than replacing them; fundamentals like verification procedures and detection engineering remain the core defense.

Knowledge Check

Click an answer to reveal the explanation.

Why do traditional phishing "tells" like poor grammar and generic greetings no longer reliably indicate a phishing attempt?

Correct answer: B. Generative AI produces fluent, personalized, contextually appropriate text on demand, so the quality gap that used to expose phishing attempts is far less reliable. Detection now has to lean on authentication, infrastructure signals, and behavioral cues instead of text quality.

What is the key difference between an evasion attack and a data poisoning attack against a machine learning model?

Correct answer: B. Evasion attacks happen at inference time against an already-trained, deployed model, crafting input that crosses its decision boundary without changing the attacker's underlying intent. Data poisoning happens earlier, during training or fine-tuning, corrupting what the model learns so its future predictions are biased or backdoored.

What is MITRE ATLAS?

Correct answer: C. ATLAS (Adversarial Threat Landscape for Artificial-Intelligence Systems) is a MITRE-maintained knowledge base structured like ATT&CK, cataloguing tactics and techniques adversaries use against AI systems, including data poisoning and evasion. It is not a scanner, a compliance framework, or a commercial feed. Chapter 6 covers it in more depth.