Reference
RAG, end to end
Retrieval-augmented generation is where an attacker’s document meets your model’s trust. RAG is spread across the book because it touches every layer - the context window that ingests it, the data layer that stores it, the method that tests it. This page pulls it into one place: what the pipeline is, where it breaks, how to hold it, and which chapter carries each piece in depth.
What RAG is
RAG lets a model answer from your documents without retraining. The pipeline is six steps: ingest, chunk, embed, store, retrieve, ground. A question is turned into a vector, the closest chunks are pulled from a vector database, and those chunks are pasted into the context before the model answers. The concrete request-to-answer trace is in Start here - the basics.
The one fact that makes it a security topic: the retrieved text lands in the same context as your instructions. The model has no reliable way to tell a trusted system instruction from a sentence in a document it just fetched. So every document that can be indexed is a candidate instruction, and the vector store holds a derived copy of whatever it was built from.
The attack surface
| Attack | What it is | Depth |
|---|---|---|
| Knowledge-base poisoning | Get a malicious instruction indexed, so the model retrieves and obeys it later | VI.4 · II.2 |
| Retrieval manipulation | Craft content to win the similarity match for a target query (the PoisonedRAG line - a few passages can control answers) | VI.4 |
| Authority spoofing (DACSI) | Non-imperative, metadata-like text impersonating a provenance or policy signal, evading imperative-injection filters | II.2 |
| Indirect injection at scale | Instructions planted in non-rendered fields of pages a crawler indexes, never seen by a human | II.2 |
| Cross-tenant leakage | A shared multi-tenant store with no role-aware retrieval surfaces another tenant’s documents | V.3 |
| Embedding inversion | Reconstruct source text from stored vectors - storing embeddings is not anonymization | V.3 |
| Membership inference | Decide whether a specific record is in the store or training set, from similarity signals | I.3 |
How you defend it
The retrieved chunk is untrusted input that happens to look authoritative. The defenses treat it that way:
- Role-aware, entitlement-checked retrieval. Re-check the requesting user’s permissions at query time, against the documents being returned - not just at ingest. This is the single control that closes cross-tenant leakage. See V.3 · The data layer.
- Spotlight and delimit retrieved text. Mark retrieved content as data, never instruction, and keep it out of the instruction channel. The mechanics are in II.2 · Prompt injection & the LLM attack surface and II.5 · Guardrails - what holds, and how to prove it.
- Provenance on every chunk. Validate and sign ingested sources; tag each chunk with where it came from, and distrust chunks whose source is untrusted.
- Secure the vector store as raw data. Encrypt vectors at rest, lock down the store, and minimize what you embed - it holds the data it was derived from (V.3 · The data layer).
- Gate the ingest. Treat the corpus as an attack surface: scan your own crawl or knowledge base for instruction-like text in non-rendered fields before it is indexed.
Where it is covered in depth
- VI.4 · AI red-team playbook - exploiting RAG pipelines and attacking embeddings, with the offensive method.
- V.3 · The data layer - the vector store, retrieval entitlement, and cross-tenant isolation.
- II.2 · Prompt injection & the LLM attack surface - indirect injection, the primitive behind KB poisoning.
- I.3 · Training data - poisoning the corpus, extraction, and membership inference.
- Attack index - the RAG rows in the wider offensive path.