Reference
Reference library
The primary-source spine behind the playbook: the papers, specs, standards, and vendor advisories its claims rest on - grouped by topic, each with a short ID the body text cites inline.
Adversarial ML, privacy & LLM canon
a1Goodfellow - FGSM - arXiv:1412.6572a2Madry - PGD · Gu - BadNets - arXiv:1706.06083 · 1708.06733cCarlini - Extracting Training Data from LLMs - USENIX Security; arXiv:2012.07805pCarlini - Poisoning Web-Scale Training Datasets is Practical - arXiv:2302.10149ppZhang - Persistent Pre-training Poisoning of LLMs - arXiv:2410.13722scScaling Trends for Data Poisoning in LLMs - arXiv:2408.02946nkPoisoning Attacks on LLMs Require a Near-constant Number of Poison Samples - arXiv:2510.07192 (Anthropic / UK AISI / Alan Turing, Oct 2025: ~250 docs backdoor 600M-13B models regardless of scale)gGreshake - Indirect Prompt Injection - arXiv:2302.12173zZou - Universal Transferable Attacks (GCG) - arXiv:2307.15043saHubinger - Sleeper Agents (Anthropic) - arXiv:2401.05566wWillison - lethal trifecta - simonwillison.netskPrompt Injection Attacks on Agentic Coding Assistants - vulnerabilities in skills, tools and protocol ecosystems - arXiv:2601.17548ncNasr, Carlini et al. - Scalable Extraction of Training Data from (Production) LLMs - arXiv:2311.17035 (the “divergence” attack that pulled verbatim training data incl. PII from ChatGPT)elEchoLeak - zero-click indirect injection in M365 Copilot (CVE-2025-32711, CVSS 9.3; “LLM Scope Violation”) - Aim Labs; patched Jun 2025
Multimodal attacks
ipiNagaraja - Image-based Prompt Injection (IPI) - arXiv:2603.03637uvUniversal Adversarial Attack on Aligned Multimodal LLMs - arXiv:2502.07987ciCSA - Image Prompt Injection in Multimodal LLMs - CSA Labs, Mar 2026mmSeeing the Threat - VLM adversarial-attack study - arXiv:2505.21967
Agent protocols (MCP / A2A)
mMCP Specification (Authorization), 2025-11-25 - OAuth 2.1 RS; RFC 9728/8707msMCPShield · MCPSecBench - arXiv:2604.05969 · 2508.13220cmComparative Threat Model - MCP/A2A/Agora/ANP - arXiv:2602.11327aA2A Protocol Specification - a2a-protocol.orgmeSecuring an A2A Application (MAESTRO) - arXiv:2504.16902sSurvey of Agent Interoperability Protocols - arXiv:2505.02279adAgentDojo - dynamic env for prompt-injection attacks/defenses on LLM agents · InjecAgent - indirect-injection benchmark for tool-using agents - arXiv:2406.13352 · 2403.02691mcparchMCP architecture & specification - primitives, transports, lifecycle - modelcontextprotocol.io (primary source)mcpsaWhen MCP Servers Attack - taxonomy, feasibility, mitigation · MCP threat modeling & tool-poisoning injection - arXiv:2509.24272 · 2603.22489u42Unit 42 - new prompt-injection vectors through MCP sampling - Palo Alto Networks, 2026mcpverMCP - protocol versioning & the current stable revision - modelcontextprotocol.io (primary source; check before citing any spec version)mcpdepMCP - deprecated features registry · Feature lifecycle policy - 12-month window, 90-day expedited floor - modelcontextprotocol.iomcpextMCP extensions - reverse-DNS identifiers, opt-in, independent versioning · Extension client support matrix - modelcontextprotocol.iomcpcliMCP client concepts - Roots, Sampling, Elicitation; roots are not a security boundary - modelcontextprotocol.iomcpcbpMCP client best practices - programmatic tool calling, progressive discovery, per-call authorization - modelcontextprotocol.iomcpchgMCP changelog 2025-11-25 · draft changelog - the SEP-to-change map - modelcontextprotocol.iomcpsdkMCP SDK tiers - conformance & security-response commitments · SDK list by tier · conformance suite - modelcontextprotocol.io / GitHubmcptaMCP Tool Annotations Interest Group - six competing SEPs open - modelcontextprotocol.io, Apr 2026ciscmcpCIS Controls v8.1 - Model Context Protocol Companion Guide v1.0 - Center for Internet Security, Apr 2026. 82pp; interprets all 18 Controls for MCP, four deployment patterns with trust-boundary diagrams, and an appendix mapping ~20 MCP CVEs to Safeguards. License CC BY-NC-ND 4.0. Companion guides also exist for AI/LLM and AI agentsowaspastOWASP Agentic Skills Top 10 (AST01-AST10) - OWASP GenAI Security Project. Working draft, in community review at time of writing - the risk set covers agent “skill” bundles (SKILL.mdinstructions plus code and resources). Cite as emerging, not settled; re-check for a published release before relying on the item numberingmcpsbpMCP Security Best Practices, 2026-07-28 revision - modelcontextprotocol.io. Named attacks with normative mitigations: confused deputy, token passthrough, SSRF, state handle hijacking, local server compromise, OAuth URL injection, stdio proxy escalation, mix-up attacks and localhost redirect impersonation, plus scope minimizationmcp2607MCP2026-07-28stable release - GitHub, 28 Jul 2026. “This release marks the stable release of the2026-07-28revision.” The versioning page still named 2025-11-25 as current on 29 Jul 2026, so cite bothmcp52869CVE-2026-52869 - MCP Python SDK routes sessions by id without verifying the principal - NVD, 15 Jul 2026. CVSS 7.1, fixed 1.27.2. The empirical case behind the spec’s State Handle Hijacking section
Browser / computer-use agents
bcAgentic AI security 2026 - infrastructure-level injection, CoSAI surface map - Adversa AI / CoSAI, May 2026
Coding agents & Codex
caOpenAI - Introducing upgrades to Codex - openai.comcsOpenAI - Codex agent approvals & security - developers.openai.comcyOpenAI - GPT-5.3-Codex system card / cyber safeguards - deploymentsafety.openai.com, Feb 2026
Offensive AI & frontier safety
xAnthropic - Disrupting AI-orchestrated espionage (GTG-1002) - Nov 2025rsAnthropic - Responsible Scaling Policy - RSP v3.4, 2026lpOpenAI - Preparedness Framework v2 · DeepMind - Frontier Safety Framework v3.1 - Apr 2025 · Apr 2026mtMETR - Common Elements of Frontier AI Safety Policies - metr.org, Mar 2025afAffordance analysis of the OpenAI Preparedness Framework (critique) - arXiv:2509.24394cotChain of Thought Monitorability: A New and Fragile Opportunity for AI Safety - arXiv:2507.11473 (multi-org: OpenAI, Anthropic, GDM, et al., Jul 2025)
Threat modeling
tmCSA - MAESTRO agentic AI threat-modeling framework (7 layers) - cloudsecurityalliance.org, 2025taSecuring Agentic AI - MAESTRO applied to a monitoring agent - arXiv:2508.10043
Singapore AI testing & accreditation
smProject Moonshot - LLM benchmarking & red-teaming toolkit - AI Verify FoundationstaSingapore AI Tester Accreditation Programme - The Edge Singapore, 2026siISO/IEC 42119-8 (proposed/draft) - benchmarking & red-teaming methodology - IMDA, Apr 2026iskIMDA - Starter Kit for Testing LLM-Based Applications - imda.gov.sg, 2025
High-harm capability evaluation
huMeasuring Harmful Capability Uplift (human-centered evals) - MIT, arXiv:2603.26676, Mar 2026hqQuantifying CBRN Risk in Frontier Models (WMDP/FORTRESS/VCT) - arXiv:2510.21133hfNova Premier under the Frontier Model Safety Framework (audited CBRN evals) - arXiv:2507.06260heEpoch AI - Do biorisk evals actually measure the risk? - epoch.ai, 2025hjLLMs Outperform Experts on Challenging Biology Benchmarks (VCT, LAB-Bench, GPQA-Bio, WMDP) - Justen, arXiv:2505.06108, 2025hsLLM Novice Uplift on Dual-Use, In Silico Biology Tasks (uplift study design) - Scale AI, arXiv:2602.23329, 2026
Jailbreaks & guardrail bypasses
jsJailbreakRadar - Assessment of Jailbreak Attacks · Guarding the Guardrails (taxonomy) - ACL 2025 · arXiv:2510.13893jbRobustness of LLM Safety Guardrails vs Adversarial Attacks - arXiv:2511.22047jxAnthropic - Many-shot Jailbreaking - anthropic.comjmMicrosoft - Skeleton Key & Crescendo - microsoft.com, Jun 2024jhHiddenLayer - Policy Puppetry universal bypass - hiddenlayer.com, Apr 2025jrRepello - AI jailbreak techniques & safeguards - repello.ai, 2026
Standards, verification & maturity
omOWASP MCP Top 10 - protocol-level risk taxonomy (beta) - owasp.org, 2025-2026svOWASP AISVS - AI Security Verification Standard - owasp.orgvssOWASP AIVSS - AI Vulnerability Scoring System (v0.8) - aivss.owasp.org, 2026aimaOWASP AI Maturity Assessment (AIMA) - owasp.orgstAI Red Teaming 2026 field guide - garak, PyRIT, methodology - 2026snCSA note - NIST COSAiS & AI agent security standards - CSA, Mar 2026grtOWASP GenAI Red Teaming Guide - model, implementation, infrastructure & runtime testing - OWASP, Jan 2025nahNIST - Strengthening AI Agent Hijacking Evaluations (the ASR discipline: adaptive attacks, repeat trials) - NIST, Jan 2025msrtMicrosoft - Lessons from Red Teaming 100 Generative AI Products - Microsoft AI Red Team, arXiv:2501.07238, Jan 2025jbkWei, Haghtalab & Steinhardt - Jailbroken: How Does LLM Safety Training Fail? (the two failure modes) - arXiv:2307.02483, 2023
Frameworks, Singapore & EU
mgIMDA - Model AI Governance Framework for Agentic AI (v1.5, 20 May 2026) - imda.gov.sg, launched WEF Davos 22 Jan 2026oOWASP Top 10 for LLM Apps (2025) · Agentic Top 10 (Dec 2025) - genai.owasp.orgsfGoogle SAIF · CoSAI · MITRE ATLAS · NIST AI RMF - frameworkssa2CSA Advisory AD-2026-004 - Frontier AI Risks - csa.gov.sg, 15 Apr 2026s2CSA Guidelines & Companion Guide · Securing Agentic AI Addendum · EU AI Act - csa.gov.sg · EU
Defenses & mitigations
slDefending Against Indirect Prompt Injection with Spotlighting - Hines et al., Microsoft, 2024ccConstitutional Classifiers - defending against universal jailbreaks - Anthropic, 2025 (arXiv 2501.18837)cbImproving Alignment and Robustness with Circuit Breakers - Zou et al., NeurIPS 2024clDefeating Prompt Injections by Design (CaMeL) - Debenedetti et al., Google DeepMind, 2025
Identity, detection & response
niOWASP - Agentic Top 10 ↔ Non-Human Identities Top 10 cross-map - genai.owasp.org, Dec 2025otOpenTelemetry - GenAI semantic conventions - opentelemetry.iosoSurvey of Agentic AI & Cybersecurity (defensive use) - arXiv:2601.05293oauthOAuth token-binding & delegation RFCs for NHI: DPoP 9449 · mTLS-bound 8705 · token exchange 8693 · introspection 7662 - IETFspiffeWorkload identity & policy-as-code: SPIFFE/SPIRE · OPA · Cedar - per-agent identity + runtime authorization
Data-layer security
doOrca Security - Exposed Vector Databases - orca.security, 2026de2026 AI Security Predictions - “Breach-by-Exhaust” - BigDATAwire, Dec 2025
ML supply chain & model-file security
pkReversingLabs - nullifAI: malicious ML models evading picklescan - reversinglabs.com, Feb 2025jfJFrog - Malicious Hugging Face models with silent backdoor - jfrog.com, 2024sftsafetensors - safe, code-free model serialization - Hugging FacemscProtect AI - ModelScan · Trail of Bits - Fickling - model-file scannerskcveCVE-2025-1550 - Keras .keras-archive (config.json) arbitrary code execution on model load - NVD (the model-file-is-code RCE class, beyond pickle)omsOpenSSF Model Signing (OMS) v1.0 · sigstore/model-transparency - OpenSSF/Google/NVIDIA, Apr 2025ssSLSA - Supply-chain Levels for Software Artifacts · Sigstore - slsa.dev · sigstore.devmbCycloneDX ML-BOM (v1.7) · OWASP SCVS - OWASP, Oct 2025mcdMitchell - Model Cards for Model Reporting - FAT* 2019; arXiv:1810.03993
MLSecOps & guardrails
lgProtect AI - LLM Guard · NVIDIA NeMo Guardrails · Guardrails AI - open-source runtime guardrailslfMeta - LlamaFirewall (PromptGuard 2, CodeShield) - ai.meta.comprPoisonedRAG - knowledge-corruption attacks on RAG - USENIX Security 2025; arXiv:2402.07867
AI threat libraries & emerging threats
atlMITRE ATLAS - matrix, Navigator & case studies (16 tactics, ~84 techniques, ~56 sub-techniques, v2026.06 June 2026; monthly calendar-versioned cadence) - atlas.mitre.org, 2026biBIML - Architectural Risk Analysis of ML / LLMs (BIML-78; 23 black-box risks) - berryvilleiml.com; IEEE Computer, Apr 2024mrMIT AI Risk Repository (1,700+ risks) - airisk.mit.eduaidAI Incident Database - incidentdatabase.aiavAVID - AI Vulnerability Database - avidml.orgw2Cohen, Bitton & Nassi - Morris II: zero-click GenAI worms - arXiv:2403.02817, 2024
MCP server hardening
mhMCP - Security Best Practices (confused deputy, no token passthrough, no session auth) - modelcontextprotocol.iocwCoSAI WS4 - Secure Design Patterns for Agentic Systems: MCP security - cosai-oasis, 2026nsamcpNSA - MCP: Security Design Considerations for AI-Driven Automation - NSA/AISC, May 2026mcpcsOWASP MCP Security Cheat Sheet - OWASP Cheat Sheet SeriesawsmcpCVE-2026-16584 - AWS API MCP Server skips policy checks after an init failure - NVD, 23 Jul 2026. CWE-455, 0.2.13 to 1.3.46, fixed 1.3.47. The 2026 reference case for a guardrail failing open for a whole process lifetimemcprbCVE-2025-66414 (TypeScript SDK, fixed 1.24.0) and CVE-2025-66416 (Python SDK, fixed 1.23.0) - NVD, both CVSS 8.1. DNS-rebinding protection present but off by default; the two SDKs sit on different version linesccwfCVE-2026-54316 - Claude Code pre-approvedhuggingface.coas a bare hostname for WebFetch - NVD, 23 Jun 2026. 0.2.54 to 2.1.162, fixed 2.1.163. A domain-level allowlist is not an exfiltration control when the domain is attacker-writable
Shadow AI discovery & governance
sdShadow AI - scale & risk (98% unsanctioned use; Netskope 223 violations/mo; CrowdStrike 2026) - Vectra AI, 2026seMicrosoft Entra - Shadow AI discovery in Global Secure Access - learn.microsoft.com, 2026spTenable - Shadow AI & AI-SPM (discover shadow AI in build before prod) - tenable.com, 2025st2Shadow-AI detection tooling landscape (Lasso, Nightfall, et al.) - Netwrix, 2026sgShadow AI governance - approved alternatives cut unsanctioned use ~89% - Forcepoint, 2026
AI governance, risk & maturity standards
irISO/IEC 23894:2023 - AI guidance on risk management - ISO/IEC; the risk process, built on ISO 31000imISO/IEC 42001:2023 - AI management system (AIMS) - ISO/IEC; PDCA + Annex A controlsngNIST AI RMF 1.0 (Govern/Map/Measure/Manage) & AI 600-1 GenAI Profile - NIST, 2023-2024nmNIST AI 100-2e - Adversarial ML: a taxonomy & terminology of attacks and mitigations - NIST (the canonical attack-naming scheme; rev. 2025)ecEC-Council Global Services - ADG (Adopt · Defend · Govern) framework - aigovernance.eccouncil.org, 2026