Use separate instances when segregation is the only reliable way to keep sensitive populations apart, but prefer document-level filtering when you need one knowledge base with fine-grained entitlements. The decision should follow the maturity of your metadata, role design, and operational tolerance for duplication.
What drives the architectural choice between separate instances and filtering?
The core decision is whether segregation is a control boundary or merely a convenience layer. Separate RAG instances reduce the blast radius when different populations must never co-mingle, while document-level filtering preserves a shared retrieval system and relies on authorization discipline at query time. The right answer depends on how cleanly you can express entitlements in metadata and how much duplication your operating model can tolerate.
Separate instances are strongest when the risk is structural, such as legal, contractual, or highly sensitive partitions that should not depend on runtime filters to stay isolated. Document-level filtering is stronger when the main problem is selective exposure within one corpus, because it lets you keep a single index and enforce access at retrieval time. A useful way to think about it is: if the entitlement model is brittle, duplicate the knowledge base; if the entitlement model is trustworthy, filter.
The practical difference is that separate instances shift complexity into data duplication, index management, and consistency control, while document-level filtering shifts complexity into metadata quality, role design, and enforcement correctness. That trade-off matters because retrieval quality, freshness, and operator workload can degrade quickly if teams pick the wrong pattern for the maturity of their content governance.
Where the control failure usually happens
Most bad outcomes come from assuming that filtering is only a search feature. If the system can retrieve a document before entitlements are checked, or if metadata is incomplete, stale, or inconsistently applied, sensitive content can leak through ranking, summarization, embeddings, or cached retrieval paths. Teams evaluating permission-aware retrieval should also review the underlying index and vector-store controls, as described in NHIMG’s Permission-Aware RAG Guide.
Separate instances fail differently. They reduce cross-population leakage, but they can create shadow copies, configuration drift, and inconsistent document lifecycles if the same source content is replicated across multiple corpora. That means the security benefit is real, but only if instance boundaries are operationally enforceable and the team can keep content synchronized without weakening governance.
Filtering also depends on the quality of the identity and access model behind it. If role design is too coarse, you end up overexposing content to make retrieval usable. If it is too granular, operators start bypassing the model or creating brittle exceptions. In practice, document-level filtering works best when the metadata schema is stable, the ownership model is clear, and the entitlement rules are already trusted elsewhere in the stack.
How mature teams decide which pattern to use
A mature team starts by asking whether one corpus can safely serve multiple populations without weakening the strongest required boundary. If the answer is yes, document-level filtering usually gives better reuse and lower duplication cost. If the answer is no, instance separation is the safer design because it removes the burden from every retrieval decision and from every future metadata correction.
- Use separate instances when a failure to filter even once would be unacceptable, or when populations differ enough that shared indexing would create governance ambiguity.
- Use document-level filtering when the same corpus must support many audiences and entitlement data is dependable enough to make retrieval-time enforcement trustworthy.
- Prefer the simplest pattern that your teams can audit consistently, not the most elegant pattern on paper.
The decision also changes with operational tolerance for duplication. If teams cannot afford multiple pipelines, multiple index refreshes, or multiple content governance processes, filtering may be the only sustainable approach. If they cannot afford a single entitlement mistake, duplication may be the better operational cost. NHIMG’s AI Supply Chain Security and AI-BOM Guide is useful here because duplicated knowledge paths also increase the need to track what content, tools, and dependencies each instance inherits.
Risk and Threat Considerations
The main risk is accidental overexposure, either through imperfect metadata or through a shared index that reveals documents to users who should never see them. The threat is not limited to intentional abuse, a weak filter, stale permissions, or a misrouted embedding lookup can all expose sensitive material during normal operation.
Failure mechanism: Shared retrieval paths rely on correct entitlement checks at the right stage of the pipeline. If filtering happens too late, or if the metadata required to filter is incomplete or inconsistent, the system can rank, cache, or summarize restricted content before access controls take effect.
Impact: Sensitive content can cross population boundaries, audit confidence drops, and the organization may need to choose between rapid containment and extensive reindexing. In regulated or high-trust environments, that can turn a search architecture decision into a compliance and incident-response problem.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
NIST SP 800-53 Rev 5 sets the technical controls, while ISO/IEC 27001:2022 defines the regulatory obligations.
| Framework | Control / Reference | Relevance |
|---|---|---|
| NIST SP 800-53 Rev 5 | AC-3 — Access Enforcement | Controls retrieval-time access decisions for protected documents. |
| AC-6 — Least Privilege | Supports limiting who can reach sensitive content through a shared corpus. | |
| AU-2 — Event Logging | Logging is needed to evidence who accessed which corpus or document. | |
| Recommendation — Enforce document access checks before retrieval and summarization. Restrict retrieval permissions to the minimum required population. Log retrieval and filter decisions for audit and investigation. | ||
| ISO/IEC 27001:2022 | A.5.15 — Access control | Access control governs whether one corpus can safely serve multiple audiences. |
| A.5.18 — Access rights | Access rights management underpins document-level entitlement fidelity. | |
| Recommendation — Define and enforce access rules for shared retrieval content. Review and maintain document access rights regularly. | ||
Practitioner Guidance
What to verify: Confirm that entitlements are expressed in a way your retrieval layer can enforce deterministically, including source labels, document tags, and inheritance rules. If you cannot explain how a document is excluded before ranking begins, treat the filter model as immature.
Decision rule: If a mistaken retrieval would be materially harmful, bias toward separate instances unless you can prove the metadata, role model, and operational controls are consistently reliable. If the harm is lower and the shared corpus materially improves usability, filtering is usually the better long-term pattern.
What practitioners underestimate: The hard part is not selecting the architecture, it is sustaining it as content changes. A filtering model that works in a clean pilot often fails later when taxonomy drift, role sprawl, and content duplication begin to outpace governance.
Practitioner takeaway: Choose separation when the boundary itself must be trusted, and choose filtering only when your entitlement model is strong enough that retrieval can safely depend on it every day.
Related resources from NHI Mgmt Group
- How should security teams decide whether JIT access is safe for non-human identities?
- How should security teams decide between native ERP controls and a separate governance platform?
- How should security teams decide between gateway-level control and container isolation for agents?
- How should security teams decide between RAG and MCP in production AI systems?