AI Security: Prompt Injection, Data Leakage and Red Teaming · Agents, tools, supply chain and RAG poisoning · lesson 10 of 17 · 14 min
RAG poisoning, vector store security and memory
The knowledge base is an attack surface
Retrieval-augmented generation trusts whatever lands in the index. If an attacker can get content into your corpus, or if your retrieval ignores who is allowed to see what, the model will faithfully use poisoned or unauthorized context. OWASP covers this under Data and Model Poisoning (LLM05:2026) and Vector and Embedding Weaknesses (LLM09:2026); the Agentic Top 10 adds memory and context poisoning (ASI06).
RAG poisoning
How content gets in: public web crawls, user uploads, customer reviews, supplier catalogs, wiki pages anyone can edit, shared drives, support tickets, emails ingested for search, and agent memory.
What poisoned content does:
- Carries indirect prompt injections that activate on retrieval.
- Plants false facts that the model repeats confidently (wrong prices, fake policies, a competitor's phone number).
- Games retrieval: text crafted to be highly similar to target queries so it outranks legitimate documents. Research such as PoisonedRAG (2024) showed that a small number of crafted passages can steer answers for targeted questions.
- Persists: once indexed, it affects every future user until removed.
Controls:
- Source trust tiers. Separate indexes or metadata for internal-authoritative, partner, user-generated and public content; weight or filter by tier per use case; show provenance in answers.
- Ingestion pipeline security. Authenticate writers; review or moderate user-generated content; scan for injection patterns, invisible characters and anomalous text; record who added what, when.
- Change monitoring. Alert on bulk changes, new documents that rank top for many queries, or near-duplicate floods.
- Answer-time checks. Groundedness evaluation against high-trust sources for critical facts (prices, policies); cite sources; abstain when only low-trust sources support a claim.
- Removal and re-index runbook. Be able to purge a document and its chunks quickly and verify it no longer retrieves.
Access control in retrieval
A frequent real-world flaw: every user's query searches the whole index, and the app filters results after retrieval, or not at all. Consequences: users see content they should not, and the model may summarize a confidential document for an unauthorized user.
Correct pattern: enforce authorization at query time, before or during retrieval.
- Store access metadata (tenant, groups, document ACL) with every chunk.
- Filter in the vector query itself using the authenticated user's attributes, or use separate indexes or namespaces per tenant for strong isolation.
- Re-check permissions on the source system for highly sensitive documents (ACLs change).
- Never let the model choose the filter values.
# retrieval with query-time authorization (pseudo-API; adapt to your vector DB's filter syntax)
def retrieve(user, query: str, k: int = 8):
flt = {
"tenant_id": user.tenant_id, # hard tenant isolation
"acl_groups": {"$in": list(user.groups)}, # document-level ACL
"trust_tier": {"$in": ["internal", "partner"]}, # use-case policy
}
hits = vector_db.search(embed(query), top_k=k, filter=flt)
return [h for h in hits if source_acl_still_allows(user, h.doc_id)] # optional re-check for sensitive docs
Embedding and vector store risks
- Embeddings are not anonymization. Research (for example Vec2Text, 2023) has shown that text can be partially reconstructed from embeddings. Protect vector stores like the source data: encryption, access control, retention and deletion.
- Multi-tenant leakage through shared indexes, caches or conversation memory.
- Deletion obligations. When a user exercises data-deletion rights under laws such as UK GDPR, their content must be removed from source systems, chunks, embeddings, caches and backups according to your retention policy.
Memory poisoning in agents
Long-term memory is a RAG index that the agent writes to itself. Attackers who can influence one conversation may plant "facts" or instructions that persist ("the user prefers payments to account X"). Controls: scope memory per user, validate and classify memory writes, never store instructions as memory, show users their memory with delete controls, and expire stale entries.
Worked example
A B2B SaaS company in Riyadh offered an assistant over each customer's documents. Their first version used one shared index with post-retrieval filtering. A penetration test found that highly specific queries could surface snippets from another tenant's documents in the model's context before the filter removed them from citations, and the model sometimes paraphrased them. The fix: per-tenant namespaces, filters in the vector query, and an automated cross-tenant test that runs in CI.
Pitfalls
- Post-retrieval filtering only.
- Treating embeddings as safe to share.
- Indexing everything without trust tiers or review.
- No purge capability.
How to measure success
Cross-tenant and unauthorized-access tests always return nothing; every indexed document has provenance and a trust tier; poisoned-document test cases are caught or neutralized; purge requests complete within your SLA.
Video lecture: RAG poisoning, vector store security and memory
Lecture coming soon · 14 chapters · about 8 minutes. Read the full transcript below.
- RAG poisoning and vector security
- Analogy: the donating library
- Poisoning
- Controls (1)
- Controls (2)
- Access control in retrieval
- Rules
- Vector store + memory
- Case: Riyadh multi-tenant assistant
- Example: two tenants, one platform
- Common mistakes
- Deeper: purge runbook
- Watch me do it: query-time authorization
- Recap
Lecture transcript
RAG poisoning and vector security
Retrieval-augmented generation trusts whatever lands in the index. If an attacker can get a document into your knowledge base, the model will quote it confidently. If your retrieval ignores who may see what, the model will summarize confidential documents for anyone who asks. In this lecture you will learn how RAG poisoning works, how to enforce access control in retrieval, and how to protect vector stores and agent memory.
Analogy: the donating library
An analogy for RAG security. Think of a library where anyone can donate books, and a librarian who answers questions by quoting whatever book seems most relevant. If someone donates a convincing fake book, the librarian will quote it. If the library has no restricted section, the librarian will happily read a confidential file to anyone who asks the right question. RAG security means checking donations, labelling which shelves are trustworthy, and locking the restricted section before the librarian can reach it.
Poisoning
How does poisoned content get in? Through public web crawls, user uploads, customer reviews, supplier catalogs, wikis anyone can edit, shared drives, support tickets, ingested emails and agent memory. Once in, it can carry indirect prompt injections that activate on retrieval, plant false facts like wrong prices or a fake phone number, or be crafted to rank highly for targeted queries. Research such as PoisonedRAG showed that a small number of crafted passages can steer answers to targeted questions. And it persists until someone removes it.
Controls (1)
Controls start with source trust tiers. Separate internal authoritative content from partner, user-generated and public content, using separate indexes or metadata, and filter or weight by tier for each use case. Show provenance in answers. Secure the ingestion pipeline: authenticate writers, moderate user content, scan for injection patterns and invisible characters, and record who added what and when.
Controls (2)
Then monitor changes: alert on bulk edits, near-duplicate floods, and new documents that suddenly rank first for many queries. At answer time, check critical facts like prices and policies for groundedness against high-trust sources, cite sources, and abstain when only low-trust sources support a claim. And have a purge runbook, so you can remove a document and all its chunks quickly and verify it no longer retrieves.
Access control in retrieval
Now access control, where a very common flaw lives. Every user's query searches the whole index, and the app filters results after retrieval, or not at all. Users then see content they should not, and the model may summarize a confidential document for an unauthorized person. The correct pattern enforces authorization at query time. Store tenant, group and document access metadata with every chunk, filter inside the vector query using the authenticated user's attributes, or use separate namespaces per tenant for strong isolation.
Rules
Two more rules. Never let the model choose the filter values; they come from the authenticated session. And for highly sensitive documents, re-check permissions on the source system, because access lists change. The lesson includes a short retrieval function showing tenant isolation, access groups and trust tiers all applied in the query.
Vector store + memory
Protect the vector store itself. Embeddings are not anonymization: research such as Vec2Text showed text can be partially reconstructed from embeddings, so protect vector stores like the source data, with encryption, access control and retention. Watch for multi-tenant leakage through shared caches and memory. And honor deletion rights under laws like UK GDPR across source systems, chunks, embeddings, caches and backups. Agent memory is simply an index the agent writes itself, so scope it per user, validate writes, never store instructions as memory, and let users view and delete it.
Case: Riyadh multi-tenant assistant
A worked example. A B2B software company in Riyadh offered an assistant over each customer's documents. Version one used a shared index with post-retrieval filtering. A penetration test found that very specific queries could pull another tenant's snippets into the model's context before the filter removed them from citations, and the model sometimes paraphrased them. The fix: per-tenant namespaces, filters inside the vector query, and an automated cross-tenant test running in CI on every change.
Example: two tenants, one platform
A simple example of query-time authorization. Two companies, a clinic and a law firm, use the same assistant platform. A user from the clinic asks for the latest policy on patient records. Wrong way: search everything, then drop the law firm's results before showing citations. The law firm's text has already reached the model. Right way: the search itself includes the clinic's tenant identifier and the user's groups as filters, taken from the login session. The law firm's documents are never even candidates. Nothing to leak, because nothing was retrieved.
Common mistakes
Common mistakes with RAG security. Filtering after retrieval and assuming the model will not use what it saw. Treating embeddings as anonymous and sharing them freely. Indexing everything from every source with no trust tiers. And having no way to purge a poisoned document quickly. Here is a question to test your setup: if you discovered a poisoned document right now, how many minutes would it take to remove it and prove it no longer appears in results?
Deeper: purge runbook
One level deeper on the purge runbook. When a poisoned document is found, the runbook deletes it at the source, deletes its chunks by document ID from the index, invalidates cached answers that cited it, and then runs the three queries that originally retrieved it, confirming none return it. Time it in a drill, so you know your real number.
Watch me do it: query-time authorization
Watch me do it with the retrieve function from the lesson, for two users of the same platform. User one works at a clinic, tenant C, in the group nurses. She asks for the patient-records retention policy. The function builds a filter from her authenticated session: tenant ID equals C, access groups include nurses, and trust tier is internal or partner. The vector search runs with that filter, so only chunks from her tenant and groups are candidates. It returns eight hits. For two of them, marked sensitive, the function re-checks the source system's access list, and one fails because her access was removed last week. So seven chunks reach the model. User two is at a law firm, tenant L. He crafts a very specific query copied from a clinic document he saw in a demo. His filter says tenant L, so clinic chunks are never candidates, however similar the embedding. He gets his own firm's documents or nothing. Now I test the anti-pattern for contrast: I run the same query without the filter and drop other tenants afterwards. The model's context briefly contained the clinic chunk. Finally, I commit a CI test that searches as tenant L for a phrase that only exists in tenant C's data and asserts zero results.
Recap
Recap. Your knowledge base is an attack surface. Tier sources, secure ingestion, monitor changes, ground critical facts in trusted sources and keep a purge runbook. Enforce access control in the query itself, protect embeddings like source data, and keep agent memory scoped and validated. Your next step: write an automated cross-tenant or unauthorized-access retrieval test and add a planted poisoned document to your red-team suite.
Key takeaways
- Anything that reaches the index can inject instructions, plant false facts or game retrieval, and it persists.
- Use trust tiers, secure and audited ingestion, change monitoring, grounded critical facts and a purge runbook.
- Enforce authorization in the vector query (or per-tenant namespaces), never after retrieval, and never from model-chosen filters.
- Protect embeddings like source data, honor deletion rights, and scope and validate agent memory.
Try it
Write an automated cross-tenant retrieval test for CI and add a planted poisoned document to your red-team suite, verifying your controls catch it.