By NHI Mgmt Group Editorial TeamBased on Descope: “Tutorial: Add ReBAC to Your RAG Pipeline With Descope” (April 23, 2026)

TL;DR: RAG systems can surface confidential documents to the wrong users because vector search ignores source-system permissions, while ReBAC preserves ownership, team membership, and sharing relationships during retrieval, according to Descope. Static roles and metadata filters do not scale once access becomes dynamic and document-level.


At a glance

What this is: This is a tutorial on adding relationship-based access control to RAG so retrieval only returns documents a user is allowed to see.

Why it matters: It matters because IAM and governance teams need authorization to survive semantic search, not sit outside it, when AI systems access internal content.


Context

RAG systems break the moment retrieval and authorisation are treated as separate problems. Once documents are embedded, the original access rules from source systems are no longer enforced by the vector layer, so semantic similarity can surface content a user was never meant to see.

The identity governance issue is not the LLM itself but the boundary between retrieval and generation. If permissions are only represented as roles or metadata snapshots, they become too rigid for document ownership, team membership, and ad hoc sharing that change constantly in enterprise environments.


Key questions

Q: What breaks when RAG pipelines rely on vector search alone for access control?

A: The pipeline can return semantically relevant but unauthorised content because similarity ranking does not know who is allowed to see each source document. Without a separate authorisation check, the model can answer from sensitive context that would be blocked in the source system, which turns retrieval into a disclosure path.

Q: Why do RBAC and metadata filters struggle with document-level AI access?

A: They assume permissions can be expressed as stable roles or prebuilt labels, but enterprise document access often depends on ownership, team membership, and ad hoc sharing. Those relationships change frequently, so hardcoded filters become brittle and expensive to maintain across thousands of documents.

Q: How should security teams govern access in RAG systems?

A: Security teams should govern RAG access at the retrieval layer, not only at authentication. That means mapping each workflow to the smallest possible set of collections, binding retrieval to user and service account entitlements, and separating sensitive corpora so one model path cannot reach unrelated business data. The goal is to limit blast radius before the model sees anything.

Q: What is the difference between RBAC and ReBAC for enterprise document access?

A: RBAC assigns access through fixed roles, while ReBAC grants access through relationships between users, teams, and documents. ReBAC is better suited to document sharing because it can express ownership and team-based access without rebuilding metadata every time the organisation changes.


Technical breakdown

Why vector search drops source-system permissions

Vector databases store chunks as embeddings for semantic similarity, not as access-controlled records. That means the retrieval layer can rank a confidential salary document alongside a public HR policy if both answer the query well. Metadata filtering helps only when permissions are simple and precomputable. In enterprise RAG, authorisation must reason over relationships, not just labels, because document access often depends on who owns the document, which team the user belongs to, and whether the item has been shared.

Practical implication: Treat retrieval as a search problem and authorisation as a separate decision layer, or sensitive content will surface before any policy check runs.

How ReBAC models document access in RAG

ReBAC represents access as relationships between users, teams, and documents. Ownership, team membership, and sharing are evaluated at request time, which makes the policy expressive enough for dynamic enterprise access patterns. Instead of storing every possible user in document metadata, the system asks whether a relationship path exists that authorises the user. That is materially different from RBAC, which assigns fixed privileges to broad roles and struggles with document-level exceptions and temporary sharing.

Practical implication: Use relationship checks when document access depends on who owns, shares, or jointly controls the content rather than on a single role.

Why post-retrieval filtering is the control point

The safest place to enforce permission in a RAG pipeline is after the vector database returns candidate chunks but before the LLM sees them. At that point, retrieval has done its job and authorisation can remove anything the user cannot access. This preserves search quality without making the vector store responsible for identity logic. If no authorised chunks remain, the system should deny access instead of letting the model answer from restricted context.

Practical implication: Put the permission decision between retrieval and generation so the model never receives unauthorised context in the first place.


Threat narrative

Attacker objective: Extract restricted enterprise knowledge through a legitimate AI query path without needing direct document access.

  1. Entry occurs when a user submits a normal-looking question to the chatbot and retrieval returns semantically relevant chunks from internal documents without checking authorisation.
  2. Credential or permission abuse follows when the pipeline passes retrieved content directly to generation, allowing access to confidential material that was never filtered against identity relationships.
  3. Impact is unauthorised disclosure of executive, financial, HR, or other restricted information through a trusted enterprise interface.

Read our 52 NHI Breaches Analysis report for a comprehensive view of breaches impacting Non-Human Identities including AI Agents.


NHI Mgmt Group analysis

ReBAC is the missing permission layer because RAG breaks the assumption that retrieval can be treated as neutral search. Once document chunks leave the source system, the original access model is no longer attached to them. The implication is that identity governance has to move into the retrieval boundary, not remain upstream in the source application.

Static RBAC is too blunt for document-centric AI access. Ownership, team membership, and sharing create relationship paths that change faster than role assignments do. That means document-level AI access cannot be governed reliably by broad job roles alone, especially in environments where a single item may be shared across functions.

Permission reasoning becomes part of the AI control plane, not a back-office compliance check. When an LLM is connected to enterprise content, the security question is whether the model is seeing authorised context, not whether the answer sounds plausible. The practical conclusion is that retrieval systems need auditable authorisation decisions at query time.

Identity data and content data must be governed together for enterprise RAG. A document chatbot is only as secure as the relationship graph behind it. If team membership, sharing, and ownership are not the source of truth for access, the AI layer will invent a permission model by omission.

Named concept: retrieval-time authorisation boundary. This is the control point between semantic search and generation where document access must be checked before any chunk reaches the model. It is the governance seam that decides whether RAG is a search feature or an exfiltration path.

From our research library:

  • Only 44% of developers are reported to follow security best practices for secrets management, exposing a significant developer behaviour gap, according to the State of Secrets in AppSec.
  • AI-related credential leaks surged 81.5% year-over-year in 2025, with the surrounding AI infrastructure leaking 5x faster than core LLM providers, according to the State of Secrets Sprawl 2026.
  • Read next: AI Agent Authorisation Guide

What this signals

Retrieval-time authorisation boundary: Enterprise RAG programmes need to treat the gap between search and generation as a control point, not an implementation detail. If that boundary is not enforced, semantic relevance can override access decisions and expose content that the source system would have protected.

Identity teams should expect more AI access patterns to depend on relationship graphs rather than on static entitlements. That shifts governance work toward auditable ownership, sharing, and team-based controls that can change as fast as the content they protect.


For practitioners

  • Enforce authorisation between retrieval and generation Check whether every RAG request evaluates document access after semantic retrieval but before the LLM receives context, so restricted chunks never enter the prompt.
  • Model access as relationships, not static metadata Represent ownership, team membership, and sharing as explicit relationships that can be evaluated at request time instead of copying user lists into document metadata.
  • Separate search from policy decisions Keep the vector database responsible for relevance ranking and the authorisation service responsible for entitlement checks, so neither layer has to do both jobs.
  • Return denial when no authorised context remains Design the application to stop at an access denied response when retrieval finds relevant content but authorisation removes everything the user may see.

Key takeaways

  • RAG can leak confidential content when semantic retrieval is allowed to outrun authorisation.
  • The tutorial’s central control is relationship-based access evaluation between retrieval and generation.
  • Enterprises that want AI over internal documents need permission checks that follow document ownership and sharing, not just roles.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

OWASP Non-Human Identity Top 10 and OWASP API Security Top 10 address the attack and risk surface, while NIST CSF 2.0 and NIST SP 800-53 Rev 5 set the governance and control requirements practitioners need to meet.

FrameworkControl / ReferenceRelevance
OWASP Non-Human Identity Top 10NHI-05 — Overprivileged NHIRAG pipelines expose content when retrieval ignores the effective access scope behind document relationships.
Recommendation — Map AI document access to NHI-05 and block any retrieval path that exceeds the user's relationship-based scope.
NIST CSF 2.0PR.AA-05 — Access Permissions, Entitlements and AuthorizationsThe article is fundamentally about enforcing authorisation before sensitive content reaches the model.
Recommendation — Apply PR.AA-05 to ensure retrieval outputs are filtered by entitlement before generation.
NIST SP 800-53 Rev 5AC-6 — Least PrivilegeOnly the minimum authorised context should be available to the LLM at generation time.
Recommendation — Enforce AC-6 so the RAG system exposes only the documents a user is entitled to access.
OWASP API Security Top 10API6 — Unrestricted Access to Sensitive Business FlowsUnauthorised RAG answers can expose sensitive business flows through a legitimate application path.
Recommendation — Treat unrestricted RAG answers as a sensitive business flow and gate them with authorisation checks.

Key terms

  • Retrieval-Time Authorization: Retrieval-time authorization is the practice of checking access rights before data is passed into an AI model or application response. It prevents overexposure by applying policy to the search or retrieval step itself, rather than relying only on post-processing or the model’s own behaviour.
  • Relationship-Based Access: An access model where entitlements are justified by the current business relationship, such as employee, contractor, student, vendor, or service account status. In practice, the relationship defines scope, duration, ownership, and review requirements.
  • Retrieval-augmented Generation: Retrieval-augmented generation is a pattern where an AI model pulls external information before generating output. The security challenge is that access rules can weaken when data is chunked, embedded, cached, or reused, so source permissions may not automatically follow the content into the model's context.
  • Permission Boundary: A permission boundary is the enforceable limit on what an identity can do, regardless of how it behaves. For AI agents, this boundary matters more than session logs because it determines whether an action is possible in the first place.

Deepen your knowledge

NHI governance, agentic AI identity, and machine identity lifecycle are core topics in our NHI Foundation Level course, the industry's only accredited NHI security programme. If you are building or maturing an IAM programme, it is worth exploring.
NHIMG Editorial Note
Published by the NHIMG editorial team on July 6, 2026.
Updated on October 6, 2026.
NHI Mgmt Group, the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org