By NHI Mgmt Group Editorial TeamBased on Authzed: “Build a Multi-Tenant RAG with Fine-Grain Authorization using Motia and SpiceDB” (December 1, 2025)

TL;DR: Building a production-style RAG pipeline with multi-tenant permissions depends on matching retrieval to relationship-based access, not just adding embeddings and vector search, according to Authzed. The identity lesson is that authorization must travel with the data path or the LLM will surface context the user should never see.


At a glance

What this is: This tutorial shows how a multi-tenant RAG pipeline pairs SpiceDB-style relationship-based authorization with vector retrieval so users only see harvest data they are allowed to access.

Why it matters: It matters because retrieval systems can leak the wrong context unless IAM and authorization checks are applied at query time, not treated as a separate wrapper around the model.


Context

Fine-grained authorization in RAG means the retrieval layer only returns context a user is entitled to see. In this article, Authzed ties that control to a multi-tenant harvest logbook, where farms, organizations, and users are linked through relationship-based permissions.

The security problem is not the LLM itself. The risk appears when embeddings and search are built without carrying the access model through chunking, retrieval, and response generation, which makes data boundaries easy to blur in a production pipeline.

For IAM teams, the useful lesson is that authorisation cannot be bolted on after retrieval. If the data path is shared across tenants, the permissions model has to follow the object relationship, the query, and the returned source material.


Key questions

Q: What breaks when retrieval permissions are too broad in RAG?

A: Broad retrieval permissions collapse data separation. A user or service account can pull executive, HR, financial, or security records into the model context even when the request does not justify it. Once that content enters the prompt window, prompt injection and response leakage become easier to trigger, and the model can become a high-speed channel for overexposed information.

Q: Why do embeddings not replace authorization in RAG pipelines?

A: Embeddings capture similarity, not entitlement. A vector store can find related text, but it cannot tell whether the user is allowed to read the source record, so authorisation must be enforced through the object relationship model.

Q: How do teams prevent tenant data leaks in multi-tenant RAG systems?

A: Carry tenant and ownership metadata from ingestion through chunking, indexing, and query-time retrieval. Then apply a permission check against the source object before the LLM sees any context, so derived artefacts cannot outlive the access rules of the original record.

Q: What is the difference between semantic search and authorized retrieval?

A: Semantic search finds the most relevant text, while authorised retrieval finds the most relevant text the caller is permitted to access. In RAG, those are separate decisions, and only the second one protects multi-tenant data boundaries.


Technical breakdown

Relationship-based access control in RAG retrieval

Relationship-based access control, or ReBAC, authorises access by evaluating how a user relates to an object, rather than by assigning a static role alone. In this design, a farm belongs to an organisation, users belong to that organisation, and permissions inherit through those relationships. That matters in RAG because the retriever can filter chunks by ownership and membership before the LLM ever sees them. Without that linkage, vector search can return semantically relevant but unauthorised content, especially in multi-tenant systems where the same index serves multiple users.

Practical implication: apply relationship checks at retrieval time so the model never receives context from unauthorised tenants.

Why embeddings do not solve authorisation

Embeddings improve semantic matching, but they do not encode who may read the underlying data. A vector database can identify similar content, yet similarity is not permission. The article’s pattern uses metadata to connect each chunk back to a farm and then uses SpiceDB to decide whether the requesting user may query or view that record. That separation is important: retrieval quality and access control are different problems, and mixing them leads to a false sense of security. The authorisation decision must happen on the object graph that owns the data, not on the embedding alone.

Practical implication: treat vector similarity as a ranking signal, not as an access decision.

Event-driven RAG pipelines need permission-aware state

In an event-driven pipeline, the system stores content, emits an embedding job, and later runs a query agent against the indexed data. That architecture makes the control problem broader than a single API check, because each asynchronous step can widen the distance between ingestion and enforcement. If the system does not preserve tenant and user context across those steps, background processing can create unauthorised retrievability even when the front door is checked. The key design rule is to keep permissions attached to the data object and its derived artefacts throughout the workflow.

Practical implication: preserve identity and tenancy metadata across every asynchronous step that touches retrieved content.


Threat narrative

Attacker objective: The objective is to elicit unauthorised tenant data through a legitimate-looking RAG query and use the model to surface it back in natural language.

  1. Entry occurs when a shared RAG workflow ingests multi-tenant content into a common retrieval pipeline without carrying permissions forward into the derived artefacts.
  2. Credential or permission abuse occurs when the query path can reach embeddings and chunks without checking the user’s relationship to the underlying farm or organisation.
  3. Impact follows when the model returns context from a tenant the user should not be able to read, turning semantic search into a data exposure path.
  • reviewdog Action compromise 2025: A stolen maintainer token poisoned reviewdog/action-setup, leaking CI secrets including the tj-actions bot token used in the next attack.
  • CI/CD pipeline exploitation case study: Credentials in an exposed .git/config let a researcher edit a Bitbucket pipeline so it planted their SSH key on the server. No victim was named.

Read and download The State of NHI & AI Agent Breach Report 2026, covering 200+ breaches impacting Non-Human Identities including AI Agents.


NHI Mgmt Group analysis

Fine-grained retrieval authorization is now part of the data plane, not a wrapper around it. RAG systems do not fail only at prompt time. They fail when retrieval is allowed to cross tenancy boundaries before the LLM is even invoked, which means the authorisation model must sit on the retrieval path itself. For practitioners, this shifts authorization from an API concern to a core part of the knowledge access architecture.

Relationship metadata is the control surface that makes multi-tenant RAG governable. In the article’s model, farms, users, and organisations are linked through inherited permissions, and that relationship graph determines what the retriever may expose. This is the right unit of control because semantic similarity alone cannot distinguish between relevant and permitted. Practitioners should treat object relationships as the source of truth for access decisions.

Chunking without entitlement context creates identity blind spots. Once documents are split into embeddings, the system is no longer handling a single record but a set of derived access-bearing artefacts. If those artefacts lose their source ownership, downstream retrieval becomes difficult to govern. The practical conclusion is that derived content must retain the identity context of the source object.

RAG governance now spans IAM, IGA, and machine-style access paths. The same programme that governs human access reviews must also govern how retrieval workflows inherit, enforce, and audit permissions. This is not a model-risk issue alone. It is an identity and access governance issue that touches lifecycle, auditability, and least-privilege enforcement across the full query chain.

What this signals

Permission-aware retrieval is becoming the defining control for enterprise RAG. As organisations move from demos to production, the main failure mode is no longer answer quality alone. It is whether the retrieval pipeline can prove that every chunk it returns belongs to the caller’s authorised scope.

Fine-grained authorization must follow derived content. Once a document becomes chunks, embeddings, and cached context, it has become a new access-bearing object. IAM teams should expect that governance will increasingly be judged by how well those derived objects retain source entitlement and audit traceability.


For practitioners

  • Enforce permissions at retrieval time Check relationship-based access before chunks are returned to the LLM, not after the answer is generated.
  • Preserve tenant metadata on derived chunks Attach farm, organisation, and user context to every embedding and chunk so access decisions can be re-evaluated downstream.
  • Separate ranking from authorisation Use vector similarity to rank content, but require a permission check against the source object before any context is exposed.
  • Audit query paths for cross-tenant leakage Trace which users can trigger retrieval, which objects they can reach, and whether background jobs can surface data outside their tenant.

Key takeaways

  • RAG pipelines create an access problem as much as a model problem when retrieval is shared across tenants without permission checks.
  • The article’s central design pattern is to bind query-time access to the underlying relationship graph, not to the embedding itself.
  • Practitioners need to treat derived chunks and retrieval traces as governed identity-bearing artefacts, or unauthorised context can leak through legitimate queries.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

OWASP Non-Human Identity Top 10 and OWASP Agentic AI Top 10 address the attack and risk surface, while NIST SP 800-53 Rev 5 and NIST CSF 2.0 set the governance and control requirements practitioners need to meet.

FrameworkControl / ReferenceRelevance
OWASP Non-Human Identity Top 10NHI-05 — Overprivileged NHIRAG pipelines can overexpose chunks when retrieval ignores entitlement boundaries.
NHI-08 — Environment IsolationMulti-tenant RAG needs tenant isolation across embeddings, chunks, and responses.
Recommendation — Bind retrieval checks to NHI-05 so derived context never exceeds the caller’s entitlement. Isolate tenant-scoped retrieval paths so shared indexes cannot cross data boundaries.
OWASP Agentic AI Top 10ASI03 — Identity & Privilege AbuseThe query path can abuse identity context if permissions are not enforced on retrieval.
Recommendation — Apply ASI03 controls so agent-driven retrieval cannot operate outside authorised scope.
NIST SP 800-53 Rev 5AC-6 — Least PrivilegeThe pipeline should only expose source context that the requester is entitled to access.
Recommendation — Enforce AC-6 to limit retrieved context to the minimum authorised scope.
NIST CSF 2.0PR.AA-05 — Access Permissions, Entitlements and AuthorizationsThis article is fundamentally about entitlement-aware access to retrieved content.
Recommendation — Use PR.AA-05 to validate entitlements before the LLM receives any source material.

Key terms

  • Relationship-Based Access: An access model where entitlements are justified by the current business relationship, such as employee, contractor, student, vendor, or service account status. In practice, the relationship defines scope, duration, ownership, and review requirements.
  • Permission-Aware Retrieval: Permission-aware retrieval is the practice of enforcing access control before data reaches the model. It ensures the model only sees documents, records, or context that the requesting identity is allowed to access. In multi-tenant or sensitive environments, this control is more important than post-processing filters.
  • Derived Access-Bearing Artefact: A piece of content created from a source record, such as a chunk, embedding, cache entry, or retrieved snippet, that still carries access implications. In governed RAG, these artefacts must inherit the source object’s identity and entitlement context or they become difficult to control.
  • Multi-Tenant RAG: A retrieval-augmented generation system that serves more than one customer, team, or business unit from shared infrastructure. The security challenge is preserving strict data separation while still allowing semantic search and natural language querying across authorised content.

Deepen your knowledge

NHI governance, agentic AI identity, and machine identity lifecycle are core topics in our NHI Foundation Level course, the industry's only accredited NHI security programme. If you are building or maturing an IAM programme, it is worth exploring.
NHIMG Editorial Note
Published by the NHIMG editorial team on June 11, 2026.
Updated on October 10, 2026.
NHI Mgmt Group, the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org