TL;DR: RAG connects models to live internal knowledge, but that architecture creates security risks around poisoned documents, over-permissioned retrieval, data leakage, embeddings, and third-party components that legacy tools were not built to govern, according to WitnessAI. The governance problem is structural: access, trust, and output controls must be designed for the retrieval pipeline, not assumed from traditional IAM or DLP.
At a glance
What this is: WitnessAI outlines why RAG creates a distinct security problem, with risks ranging from poisoned documents and over-broad retrieval to data leakage and supply chain exposure.
Why it matters: IAM, IGA, and security teams need to govern the retrieval layer, not just the model, because RAG turns enterprise knowledge access into a live authorisation and data protection problem.
Context
Retrieval-augmented generation connects large language models to live enterprise knowledge, which means the security problem is no longer limited to the model itself. The primary governance issue is that documents, retrieval scopes, embeddings, and outputs all become part of the trust boundary, but many existing controls were built for static applications or conventional data access patterns.
That mismatch matters because RAG systems can surface financial records, customer data, internal policies, and proprietary research at query time. When access, provenance, and output handling are not designed for that pipeline, organisations inherit a risk that looks like IAM, DLP, and application security at once, but is not fully covered by any one of them.
Key questions
Q: What breaks when RAG systems rely on legacy IAM alone?
A: Legacy IAM can control who reaches an application, but it does not fully govern which documents enter the context window, how retrieval scopes are partitioned, or how model outputs are filtered. In RAG, those missing controls can turn routine access into data exposure or manipulated answers, so the retrieval pipeline needs its own policy layer.
Q: Why do poisoned documents create such a serious RAG risk?
A: Because the attacker is influencing both retrieval and generation at the same time. A poisoned document can be selected because it matches a query, then alter the model’s behaviour once it is included in context. That is why provenance, content sanitisation, and instruction hierarchy are essential controls rather than optional hardening.
Q: How should teams reduce data leakage in RAG outputs?
A: They should control leakage before the model responds by redacting or tokenising sensitive values, enforcing response authorisation, and scanning both prompts and outputs for unsafe disclosures. Output filtering alone is not enough if the retrieval layer can already assemble sensitive context for the model to repeat.
Q: What should security teams check before using RAG in incident response?
A: Check that the retrieval set includes current logs, approved playbooks, and reliable historical context, and that those sources are authenticated and isolated. If the knowledge base can be polluted or crossed between tenants, the agent may recommend the wrong containment step with high confidence.
Technical breakdown
Why RAG trust boundaries are different from conventional IAM
RAG does not simply authenticate a user and then return a fixed application response. The model receives retrieved content as context, so trust is distributed across ingestion, indexing, retrieval, and generation. That creates a pipeline where a document can influence both what is surfaced and how the model answers. Conventional IAM can tell you who may reach a system, but it does not by itself define which documents are safe to retrieve, how much context can be passed forward, or when retrieved material becomes part of the model’s decision path.
Practical implication: govern the retrieval pipeline as a distinct control plane, not as an ordinary application back end.
How poisoned documents and indirect prompt injection work
Poisoned documents are malicious content placed into a knowledge base so they are retrieved for specific queries and carry instructions that steer model behaviour. Indirect prompt injection extends the same idea through email, web pages, PDFs, or database records that the RAG system later ingests as trusted context. The LLM does not reliably distinguish human-intended data from adversarial instructions once both appear in the same context window. That is why source provenance, content sanitisation, and instruction hierarchy matter as much as retrieval accuracy.
Practical implication: verify document provenance and strip untrusted instructions before content reaches retrieval or generation.
Why vectors, embeddings, and third-party components expand the attack surface
Vector databases and embeddings are often treated as neutral infrastructure, but they carry their own security risks. Embeddings may leak information through inversion or attribute inference, while vector stores can be poisoned to influence retrieval frequency and model output. Third-party orchestration frameworks, connectors, proxy layers, and MCP servers add supply chain risk because each dependency can alter data flow or execution behaviour. In other words, the RAG stack is a chain of trust, and any weak link can affect confidentiality, integrity, and response quality.
Practical implication: extend supply chain review, logging, and access control to every retrieval dependency, not just the model provider.
Breaches seen in the wild
- Vercel Context.ai OAuth Supply Chain Breach: Shadow AI app Context.ai OAuth integration exposes Vercel customer data via unmanaged third-party token.
- Palo Alto Networks Salesforce data theft 2025: Stolen Drift OAuth tokens exposed Palo Alto Networks CRM data, including support notes where some customers had shared credentials.
Read and download The State of NHI & AI Agent Breach Report 2026, covering 200+ breaches impacting Non-Human Identities including AI Agents.
NHI Mgmt Group analysis
RAG security exposes an identity gap, not just a data-handling gap: the control problem begins when enterprise knowledge becomes queryable context. Traditional IAM assumes a user is authorised to reach a system or collection, but RAG can turn that access into downstream model exposure that the original policy never explicitly contemplated. The practitioner implication is that retrieval authorisation has to be treated as a first-class identity decision, not a side effect of application access.
Prompt injection is the clearest sign that model trust and data trust are being conflated: a retrieved document can carry instructions that the model will follow because the architecture treats context as semi-trusted input. That is why document provenance, instruction hierarchy, and sanitisation belong in the control design, not in the after-incident clean-up. The implication is that content entering the retrieval pipeline must be governed as if it can change behaviour, because in practice it can.
Over-permissioned retrieval creates identity blast radius inside the knowledge layer: if every user or agent shares the same retrieval scope, the model can expose material far outside the caller’s role. This is not a simple permission bug but a governance failure in how the knowledge base is partitioned and queried. The implication is that retrieval namespaces, collection boundaries, and query-time policy must match the sensitivity of the underlying content.
Legacy DLP is too late in the pipeline for RAG: by the time a sensitive string is detected in a response, the trust decision has already happened upstream. RAG needs controls that act before content is indexed, before it is retrieved, and before it reaches the model context window. The implication is that runtime protection has to be paired with ingestion and retrieval governance, or leakage will remain a structural possibility.
Third-party RAG dependencies are now part of the identity perimeter: orchestration frameworks, embeddings services, connectors, and MCP servers can all shape what data is trusted and how it moves. That broadens the meaning of access governance beyond accounts and tokens to include components that mediate retrieval itself. The implication is that security teams must review dependency trust the same way they review privileged access paths.
What this signals
Retrieval scope is now a governance boundary: if a model can only see the knowledge assigned to its query path, exposure becomes a policy problem rather than a broad data lake problem. Teams should expect RAG programmes to converge on finer-grained namespaces, collection-level authorisation, and stronger review of what counts as retrievable content.
Prompt injection changes the meaning of untrusted input: in RAG, external content is not merely data to display but potential instruction material that can alter behaviour. That means security teams need ingestion rules that treat web pages, emails, PDFs, and indexed records as executable influence unless proven otherwise.
RAG security will increasingly be measured at the pipeline level: organisations will judge success by whether retrieval, output handling, and dependency trust are governed together. That is the operational shift hidden in the architecture, and it will determine which AI deployments move from pilot to production.
For practitioners
- Define retrieval-level access boundaries Split knowledge sources into namespaces or collections that map to role, function, or data sensitivity, and ensure query-time policy enforces those boundaries before any context is assembled.
- Treat knowledge ingestion as privileged Require provenance checks, signed source validation, and controlled write access before any document can enter a production vector store or retrieval index.
- Sanitise untrusted content before retrieval Strip hidden instructions, normalise text, and break external content into fixed-size passages so malicious payloads cannot ride into the context window.
- Add runtime protection for prompts and responses Use bidirectional scanning, tokenisation, and response authorisation so sensitive data is reduced before the model can echo or exfiltrate it.
- Review third-party retrieval dependencies Inventory connectors, orchestration frameworks, vector databases, and agent-facing services, then validate who can change data flow, credentials, and retrieval scope.
Key takeaways
- RAG creates a distinct security problem because the model is only one part of a larger trust pipeline that also includes documents, retrieval scopes, embeddings, and third-party components.
- The main failure mode is over-trusted retrieval, where malicious or over-permissioned content can shape what the model sees and says.
- Effective governance depends on provenance checks, retrieval boundaries, sanitisation, and runtime protection working together before sensitive content reaches the model.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Non-Human Identity Top 10 and MITRE ATT&CK address the attack and risk surface, while NIST SP 800-53 Rev 5 and NIST CSF 2.0 set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP Non-Human Identity Top 10 | NHI-03 — Vulnerable Third-Party NHI | RAG depends on third-party components and connectors that shape access and trust. |
| NHI-05 — Overprivileged NHI | The article centres on retrieval scopes that expose too much internal knowledge. | |
| NHI-04 — Insecure Authentication | RAG security relies on query-time trust decisions and enforced access boundaries. | |
| Recommendation — Review third-party retrieval components for inherited trust and restrict their ability to alter data flow. Limit retrieval permissions to the smallest data scope that the calling identity actually needs. Authenticate retrieval access separately from model access and tie it to data sensitivity. | ||
| NIST SP 800-53 Rev 5 | IA-9 — Identification and Authentication (Non-Organizational Users) | RAG pipelines often rely on service and machine identities across data sources and connectors. |
| Recommendation — Apply machine-to-machine authentication controls to every retrieval and connector path. | ||
| NIST CSF 2.0 | PR.AA-05 — Access Permissions, Entitlements and Authorizations | The article is fundamentally about who can retrieve what content and when. |
| Recommendation — Align retrieval entitlements with content sensitivity and verify authorisations at query time. | ||
| MITRE ATT&CK | TA0006;TA0010 — Credential Access; Exfiltration | The source describes prompt injection, leakage, and exfiltration paths that map to credential and data abuse. |
| Recommendation — Map RAG abuse paths to credential access and exfiltration techniques to improve detection coverage. | ||
Key terms
- Retrieval-augmented Generation: Retrieval-augmented generation is a pattern where an AI model pulls external information before generating output. The security challenge is that access rules can weaken when data is chunked, embedded, cached, or reused, so source permissions may not automatically follow the content into the model's context.
- Prompt Injection (Agentic): An attack where malicious instructions are embedded in content that an AI agent reads, causing the agent to execute unintended actions using its own legitimate credentials. A primary vector for agent goal hijacking and identity abuse.
- Vector Database: A data store that indexes embeddings so semantically similar content can be retrieved quickly. In RAG systems, the vector database is part of the trust boundary because it controls what context is surfaced, how often, and under which permissions.
- Context Window: The context window is the text a model receives at one time, including prompts, retrieved documents, and conversation history. Security teams care about it because it becomes the practical boundary between trusted instructions and untrusted content, especially when the application assembles that text automatically.
Deepen your knowledge
NHI governance, agentic AI identity, and machine identity lifecycle are core topics in our NHI Foundation Level course, the industry's only accredited NHI security programme. If you are building or maturing an IAM programme, it is worth exploring.
Published by the NHIMG editorial team on June 6, 2026.
Updated on October 8, 2026.
NHI Mgmt Group, the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org