Offensive security · for the agentic age
The AI Security Playbook
How AI systems actually get attacked - and what actually stops it. Written for the people who have to test one, defend one, or sign one off.
43 chapters~84k words19 CVEs mapped one author, Iaroslav Mezin
The 90-second triage
An AI system is exploitable for data theft when it has all three at once. Break any one leg and the path closes.
Private data
Can it reach data you would not publish?
Break it scope access on-behalf-of the user, just in time
Untrusted content
Does it read text an attacker can influence?
Break it quarantine or spotlight untrusted input
External comms
Can it send data out - mail, webhook, API?
Break it allowlist egress; gate irreversible actions
The rest of the one-page cheat sheet - the trust-boundary principle, the agentic stack, where the controls live. Made to be screenshotted.
What do you have to do?
Nobody arrives here to read a book. Pick the job; each one lands on the chapter that does it.
Or take the whole path
- IThe model7 chaptersThe weights: training, poisoning, adversarial ML, the model supply chain
- IIThe context window5 chaptersThe context window: prompt injection, jailbreaks, multimodal, guardrails
- IIIThe agent loop5 chaptersThe agent loop: tools, memory, persistence, containment
- IVThe protocol layer6 chaptersThe protocol layer: model APIs, MCP, A2A, non-human identity
- VThe infrastructure3 chaptersThe infrastructure: cloud, red-teaming, the data layer
- VIThe method6 chaptersThe method: threat modeling, red teaming, running an engagement
- VIIThe program5 chaptersThe program: secure SDLC, Agent DLC, detection and IR, retirement
- VIIIGovern it6 chaptersGovern it: SAIF, NIST AI RMF, ISO 42001, jurisdictions, the advisor's playbook
Currently tracking: July 2026 - MCP 2026-07-28, and six chapters that were stubswhat changedthe incident board