CHAPTER 06 35 MIN READ ADVANCED

Securing AI and ML Systems

Chapter 3 looked at prompt injection against LLM applications specifically. This chapter zooms out. Your organization's models, training pipelines, and the infrastructure that serves predictions to users are production systems, and they need the same rigor you already apply to any other production system: access control, supply chain discipline, change tracking, and a threat model. This chapter covers model theft, AI-specific supply chain risk, MLOps security, the risks that come with retrieval-augmented generation, and MITRE ATLAS as the framework that ties it together.

model securityMLOps securityRAG and vector database riskAI supply chainMITRE ATLAS
Before you start: This chapter assumes the model-as-attack-surface framing from chapter 2 (data poisoning, adversarial examples) and the LLM application risks from chapter 3 (prompt injection). Model theft and extraction is new material, introduced in this chapter's first section. Here we treat the AI system as infrastructure your team builds, deploys, and operates, and ask how you secure it the way you'd secure any other production service.

Model Theft and Extraction

A deployed model is usually reachable through an API, and nothing about that input/output exchange is inherently secret. Model extraction (model stealing) abuses exactly that: an attacker sends a large volume of queries, collects the input/output pairs, and trains a substitute model that approximates the original's decision boundaries, without ever touching the original weights, training data, or infrastructure.

Why it mattersDetail
IP theftA production model represents months of data collection, labeling, compute spend, and tuning. Approximating it for the cost of API queries steals that investment.
Enables further attacksA local substitute can be probed offline, at unlimited scale, with no query budget or logging on your side, making it far easier to find adversarial examples that often transfer back to the original. Extraction is frequently step one, not the whole attack.
Practical mitigations, and their limits:
  • Rate limiting slows extraction but a patient attacker with multiple accounts/IP ranges can spread queries below any reasonable threshold
  • Query-pattern monitoring flags suspicious accounts, but an attacker who understands the detection logic can shape queries to look legitimate
  • Returning only the top label instead of confidence scores/logits narrows the signal an attacker gets per query, at some cost to legitimate API consumers

None of these are complete on their own, they raise the cost and query volume needed, but nothing makes extraction impossible against a public or semi-public API. Necessary, not sufficient.

AI Supply Chain Security

Security teams already know how to think about software supply chain risk: don't pull from untrusted sources, pin versions, scan for known vulnerabilities, verify provenance. AI systems introduce the same category of risk across three additional surfaces that don't map cleanly onto traditional dependency scanning.

SurfaceThe risk
Pretrained models & fine-tuned checkpointsA model file isn't source code you can read, it's numeric weights with no straightforward way to inspect for malicious behavior. A backdoored model performs normally on nearly all inputs and only misbehaves on attacker-chosen triggers, so routine evaluation against your test set can pass cleanly while the backdoor sits untriggered. Provenance is often the only signal you have before deployment.
ML framework & library dependenciesSame class of risk as any open-source supply chain, but with less mature scanning/pinning tooling for ML-specific artifacts (model files, dataset packages, weights repos) than for general-purpose package ecosystems.
Training data supply chainDatasets scraped or sourced from third parties can carry poisoned or mislabeled samples unnoticed, connecting directly to the data poisoning technique from chapter 2. A dataset is a supply chain artifact in the same sense a dependency is.

This isn't a new discipline, it's the same supply chain discipline already applied to open-source software, extended to treat models and datasets as first-class artifacts needing the same provenance checks, skepticism about unverified sources, and inventory tracking.

MLOps Security

MLOps is the set of practices and tooling for building, deploying, and operating ML models in production, roughly the same relationship to a data science team that DevOps has to a software engineering team. Wherever it exists, it's part of your attack surface and needs controls mirroring any production deployment pipeline.

Control pointWhat to check
Training pipeline & data storageWeak access control on training data means anyone who can write to that location can poison what the next run learns from. The pipeline needs the same protection as a CI/CD pipeline, because in effect it is one.
Production push accessA registry that lets any engineer promote a checkpoint straight to serving with no review step is a deploy pipeline with no gate. Same question as any code deployment: who can ship, what has to happen first, can you prove after the fact who shipped what.
Versioning & audit trailsIf a bad or backdoored model gets deployed, you need to trace which version is live, when it changed, who approved it, and what the last known-good version was to roll back quickly. Without that trail, IR against a compromised model becomes guesswork.
The pattern you already know:
  • Detection-as-code, covered in the Detection Engineering module, applies version control, peer review, and an audit trail to detection rules so a bad rule change is traceable and reversible.
  • MLOps security is the same discipline applied to models instead of rules: version control over training code and model artifacts, peer review before a model reaches production, and an audit trail tying every deployed model back to the pipeline run and reviewer that produced it.
  • If your team already runs detection-as-code, you have the mental model for this. The artifact under version control changed; the discipline around it didn't.

Last, secure the feature store and the data pipelines that feed a live model. A feature store computes and serves the input features a model consumes at inference time, often pulling from multiple upstream systems. If those pipelines can be tampered with, an attacker doesn't need to touch the model at all; corrupting the features a healthy model consumes can produce the same bad outcome as corrupting the model itself, and it's a path that's easy to overlook because it sits outside the model artifact entirely.

RAG and Vector Database Risks

Retrieval-Augmented Generation (RAG) grounds an LLM's response in an organization's own content without retraining the model: the system embeds the user's question, searches a vector database for similar documents, and inserts the best matches into the model's context alongside the question. It's the pattern behind most "chat with your company's documents" tools, popular because it avoids re-fine-tuning every time documents change. The risk is easy to miss because it doesn't live in the model, it lives in the retrieval layer.

RiskWhere it originatesWhat it connects to
Access control bypassOriginal documents had permissions (some files restricted to certain teams); if the vector database doesn't carry those forward, any user who can query RAG can potentially retrieve content they were never authorized to read directly.Data governance gaps carried into a new system
Poisoned retrieval corpusAn attacker gets a malicious document into the indexed content (shared knowledge base, monitored upload folder, compromised upstream source); once retrieved, its content is inserted directly into the model's context, and the model can't reliably distinguish reference material from instructions.Indirect prompt injection (chapter 3)
Embedding inversionEmbeddings are meant to be a lossy numeric representation, not the text itself, but research shows meaningful information can sometimes be reconstructed from them.Vector database exposure or compromise

The practical takeaway: a vector database holds a derivative of sensitive content and needs the same access controls, monitoring, and input validation as the original documents, not a lighter-touch version because it "just" holds embeddings.

MITRE ATLAS as the Reference Framework, and Where to Go from Here

Chapter 2 introduced MITRE ATLAS as the ATT&CK-style knowledge base for adversary behavior against AI systems. It's worth returning to here because everything in this chapter, model extraction, supply chain compromise, MLOps pipeline attacks, and RAG-specific risks, maps onto tactics and techniques that ATLAS already catalogs in a structured way.

ATLAS spans the full ML lifecycle the same way ATT&CK spans the enterprise kill chain: reconnaissance and resource development against an ML system, initial access, ML model access, execution, persistence, and on through exfiltration and impact. Model theft, the training data poisoning covered in chapter 2, and the supply chain and RAG risks covered in this chapter all have a home somewhere in that structure. ATLAS spans around 16 tactics and roughly 80-plus techniques at this point, and each technique entry typically includes real-world case studies and suggested mitigations, which makes it useful for more than reference reading; it's a working checklist for threat-modeling a specific AI system.

1
Map your architecture
Identify which ATLAS tactics apply to your specific AI system: a RAG chatbot has a different exposure than an internally-trained fraud model.
→
2
Apply supply chain and MLOps discipline
Verify model and dataset provenance, gate production deployments, and keep an audit trail, as covered earlier in this chapter.
→
3
Treat it as ongoing
ATLAS is updated regularly as new techniques are documented. A one-time review goes stale; AI system security needs the same continuous attention as any other production security discipline.

None of this replaces the fundamentals covered earlier in this chapter. ATLAS gives you the vocabulary and the structure to have the conversation; the actual work is still rate limiting the inference API, verifying where a model came from, gating who can push to production, and locking down the vector database the same way you'd lock down the documents it was built from.

Key Takeaways

  • Model extraction lets an attacker train a substitute model from query traffic alone, stealing IP and enabling offline discovery of adversarial examples; rate limiting and query-pattern monitoring raise the cost but don't eliminate the risk.
  • AI supply chain risk spans pretrained models and checkpoints from public hubs, ML framework dependencies, and third-party training data; treat models and datasets as first-class supply chain artifacts, not exceptions to existing supply chain discipline.
  • MLOps security mirrors detection-as-code: version control, access-gated deployment, and audit trails, applied to training pipelines and model registries instead of detection rules.
  • RAG can bypass document-level access controls if the vector database doesn't carry forward the same permissions, and a poisoned retrieval corpus is a direct path to indirect prompt injection.
  • MITRE ATLAS provides a structured, ATT&CK-style framework spanning the full ML lifecycle, with case studies and mitigations, and it's updated as the technique landscape grows.

Knowledge Check

Click an answer to reveal the explanation.

An attacker sends thousands of varied queries to a company's public inference API and uses the resulting input/output pairs to train their own model that mimics the original's behavior. What is this attack, and why does it matter beyond the loss of API traffic?

This describes model extraction (model stealing). The immediate loss is the stolen IP and training investment, but the follow-on risk is often worse: a locally-held substitute model gives the attacker an unlimited, unlogged sandbox to hunt for adversarial examples and blind spots that frequently transfer back to the original deployed model.

A team builds a RAG system so employees can ask questions about internal documents. Some of those documents were originally restricted to a specific team. What is the specific risk this scenario raises?

RAG retrieves from whatever the vector database indexes and exposes it in the model's context. If that retrieval layer wasn't built with the same access controls as the source documents, it becomes a new, often overlooked path around those controls, letting any authorized RAG user pull in content from documents they couldn't otherwise access.

What is MITRE ATLAS, and what is it best used for?

MITRE ATLAS is a knowledge base, not a scanning tool or a certification. It catalogs adversary tactics and techniques against AI systems, spanning the ML lifecycle from reconnaissance through impact, with case studies and mitigations, giving defenders a structured way to threat-model AI systems the way ATT&CK is used for traditional enterprise environments.