AI-Powered Threats
Generative AI did not invent phishing, fraud, or malware. What it changed is the cost of doing each of them well. A convincing pretext used to take a skilled social engineer real time to write; a cloned voice used to require studio equipment and a long audio sample; a malware variant that dodges signatures used to need a developer who understood the detection logic. This chapter walks through where AI is measurably lowering that cost for attackers, and where the capability is real versus where it is still speculation.
LLM-generated phishing and social engineering
Security awareness training used to lean on a reliable set of tells: stilted grammar, generic greetings, phrasing that read like it had been translated twice. A capable LLM removes most of that signal, producing fluent, contextually appropriate prose in seconds.
| Old attack pattern | What AI changes |
|---|---|
| Classic BEC (reused templates: urgent wire transfer, gift card request, fake invoice) | An LLM drafts a pretext tailored to the specific sender/recipient relationship, matching tone and register per target. The template itself stops being a reliable detection surface. |
| Spear-phishing recon (manually researching a target's role, projects, vendors, writing style) | AI tools summarize public profiles and scraped content into a usable target profile far faster than manual OSINT, making spear-phishing-grade targeting cheap enough to scale beyond high-value targets. |
Neither is a new capability, BEC and OSINT-driven spear phishing already existed. What changes is the economics: effort that used to justify itself only for high-value targets now scales cheaply to many more targets.
- Awkward grammar or phrasing as a sole indicator, since fluent generated text passes this check easily
- Generic greetings, since personalization from scraped profile data is now cheap to produce
- Obviously templated wording, since each message can be uniquely generated rather than copy-pasted
- Mismatched tone for the claimed sender, since a model can be prompted to imitate a specific register or prior writing sample
| What still works as a detection signal | Why it survives AI-generated text |
|---|---|
| Sender authentication (SPF, DKIM, DMARC) | Text quality does not change the infrastructure the message actually came from |
| Out-of-band verification for payment or credential requests | A phone call or a second channel confirms intent independent of how well the email reads |
| Behavioral and metadata anomalies | Unusual sending patterns, reply-to mismatches, and link destinations do not improve just because the prose does |
| Reporting culture and user skepticism toward urgency and pressure | Fluency does not remove the underlying social engineering pressure tactic, which is still detectable |
Voice cloning and deepfakes
Voice cloning generates new audio in someone's voice from a sample of their speech. Modern tools need very little sample audio, and executives who speak on earnings calls or company videos generate that sample just by doing their jobs.
| Mechanism | Attack pattern it enables |
|---|---|
| Voice cloning from a short public sample | An evolution of CEO fraud/vishing: a cloned voice purporting to be an executive or vendor contact creates urgency around a wire transfer, credential reset, or policy exception. A short, urgent voicemail is often enough, it doesn't need to fool a careful listener on a long call. |
| Deepfake video/image generation | Synthetic video of someone saying or doing something they didn't, or fabricated images supporting disinformation, fake IDs, or a social engineering chain. Quality varies; automated detection is an active research area with no fully reliable solution yet. |
The psychological lever is identical to older phone-based social engineering; what changes is that "I recognized the voice" stops being a safe verification method on its own. The defensive answer is process, not better ear training: any sensitive or urgent request should go through verification that doesn't rely on the same channel or claimed identity as the request itself.
AI-assisted malware and offensive tooling
There's a real and growing pattern of attackers using LLMs as a coding assistant: drafting boilerplate, explaining unfamiliar APIs, converting proof-of-concept code between languages, or helping obfuscate a payload. None of this requires the model to be malicious by design, just a competent assistant that doesn't refuse the request.
| Claim | Established? | Why |
|---|---|---|
| AI lowers the skill floor for less sophisticated actors | Yes | Guardrails on mainstream LLMs push misuse toward jailbreaks and open-weight models fine-tuned to remove restrictions, marketed as "uncensored" coding assistants in criminal forums. |
| AI speeds up malware variant generation | Yes | A model can rapidly produce functionally equivalent versions of the same malicious code with different structure/naming/packing, each evading a rule tuned to the previous sample. Not new in concept (polymorphism is old), just cheaper. |
| LLMs autonomously discover and weaponize novel zero-days without human direction | No | Not a demonstrated, reproducible capability as of this writing. Treat vendor/commentary claims to the contrary skeptically. |
The honest framing: AI is a capable assistant that speeds up and broadens access to offensive tooling development. It is not, currently, an autonomous threat actor.
Adversarial machine learning: evasion and data poisoning
So far this chapter has covered attackers using AI as a tool. This section covers something different: attacks that target machine learning models themselves, including the ones defenders rely on for malware classification, spam filtering, and fraud/anomaly detection.
| Attack type | When it happens | What it targets | How it works |
|---|---|---|---|
| Evasion / adversarial example | Inference time, against a deployed model | Causes a specific malicious input to be misclassified as benign | Modifies an input (e.g. a malicious file) so it crosses the model's decision boundary without changing its actual behavior. The attacker doesn't need to know the model's internals, many techniques work by probing outputs and iterating, or transferring adversarial inputs crafted against a similar model. |
| Data poisoning | Training or fine-tuning time | Corrupts what the model learns, biasing future predictions or planting a backdoor trigger | Influences the data a model learns from (a public dataset, scraped content, user-submitted retraining samples) without ever breaching the model itself. Especially relevant when fine-tuning on user-supplied or less-trusted data. |
This is a good point to introduce MITRE ATLAS, the Adversarial Threat Landscape for Artificial-Intelligence Systems, a MITRE-maintained knowledge base of real-world adversary tactics and techniques against AI systems, built in the same style as ATT&CK.
- Structured like ATT&CK: a matrix of tactics, with techniques and sub-techniques mapped under each, plus case studies and mitigations
- As of its most recent 2025 update, covers around 16 tactics and roughly 80-plus techniques, including data poisoning, evasion, model extraction, and prompt-based attacks
- Exists because ATT&CK was built for conventional IT infrastructure and doesn't natively capture attacks on models, training pipelines, or AI-specific components
- Covered in more depth in Chapter 6, once the focus shifts to defending AI systems directly
What this means for defenders
Nothing in this chapter replaces the existing threat landscape, it sits on top of it. None of these AI-enhanced techniques change the underlying goal an attacker is trying to reach; what changes is the economics, cheaper, faster, more convincing at scale, putting attacks that used to require patient, well-resourced actors within reach of many more of them.
- Verification procedures that assume any request could be fabricated
- Detection logic built on behavior and infrastructure, not surface text/audio/video polish
- Awareness training that teaches people to recognize pressure tactics, not typos
Chapter 3 turns to a related but distinct problem: not the attacker using AI as a tool against your defenses, but the target being your own organization's AI or LLM application, through prompt injection and the broader risks of deploying LLM-powered features. Keep the two mentally separate going forward.
Key Takeaways
- Generative AI removes the classic phishing tells, fluency, personalization, and correct grammar no longer indicate a legitimate sender, so detection has to rely on authentication, infrastructure, and behavior instead of text quality.
- AI-assisted OSINT makes spear-phishing-grade targeting cheap enough to apply broadly, not just to high-value targets.
- Voice cloning and deepfakes undermine identity verification based on recognizing a voice or a face; the fix is process-based out-of-band verification, not better ear training.
- AI-assisted malware development mostly lowers the skill floor and speeds up variant generation; claims of autonomous zero-day discovery by LLMs are not an established capability.
- Evasion attacks manipulate inputs at inference time to fool a deployed model; data poisoning corrupts what a model learns during training. They are different attacks against different points in the ML lifecycle.
- MITRE ATLAS is a knowledge base of adversary tactics and techniques against AI systems, modeled on ATT&CK, covering roughly 16 tactics and 80-plus techniques as of its 2025 update. Chapter 6 goes deeper on defending against it.
- AI changes the economics of existing attacks rather than replacing them; fundamentals like verification procedures and detection engineering remain the core defense.
Knowledge Check
Click an answer to reveal the explanation.
Why do traditional phishing "tells" like poor grammar and generic greetings no longer reliably indicate a phishing attempt?
What is the key difference between an evasion attack and a data poisoning attack against a machine learning model?
What is MITRE ATLAS?