Security teams should prevent leakage by design, not by hoping the model behaves. Limit the model’s data scope to what the invoking user already can access, block external data sources, and require explicit user acknowledgement before any resource modification. This reduces prompt injection impact, prevents jailbreak-driven exfiltration, and keeps the agent inside a tightly governed permission boundary.
Limit the Model to the User’s Actual Data Boundary
The first control is to treat the model as a governed interface, not a trusted analyst. If the invoking user cannot already see a record, message, attachment, or internal note, the model should not be able to surface it either. That means aligning retrieval, context assembly, and response filtering to the same access rules that protect the underlying system, especially when data arrives from multiple stores or connected tools.
For teams building on conversational workflows, this is why least privilege matters more than prompt quality. A model that is technically capable of summarising sensitive content is still unsafe if the surrounding application over-broadens its view, because the user then receives data through a new path rather than through an existing authorised one.
This boundary discipline is reinforced by OWASP Non-Human Identity Top 10 and by the broader access-control expectations in NIST SP 800-53 Rev 5 Security and Privacy Controls when the model is acting through service credentials or backend APIs.
Reduce Exfiltration Paths Before You Add Model Capability
Leakage is rarely caused by the model alone. It usually happens when the system gives the model too many places to look, too much authority to act, or too much freedom to pass data onward. Blocking unnecessary external data sources, constraining tool use, and removing broad connectors sharply reduces the chance that prompt injection, indirect prompt injection, or a malicious instruction will turn the assistant into a data relay.
Teams should also treat “external sharing” as a policy decision, not a default convenience. If the assistant can call outside services, send content to third parties, or write into downstream systems, the security question is not only whether the model can answer the user, but whether it can move sensitive material outside the original trust boundary.
That design pattern aligns well with NIST AI Risk Management Framework, NIST SP 800-207 Zero Trust Architecture, and the API and tool-access risks reflected in OWASP Agentic AI Top 10.
Make High-Risk Actions Explicit, Logged, and User-Confirmed
Resource modification should not be an implied side effect of conversation. When the assistant is allowed to create, delete, transfer, publish, or share anything, require an explicit user acknowledgement step that makes the action and target visible before execution. This reduces silent misuse, makes social engineering harder, and gives security teams a point where they can inspect intent rather than infer it after the fact.
The same logic applies to responses that may reveal confidential data indirectly, such as summarising ticket history, drafting outbound messages, or composing records from mixed sources. If the action can alter state or expose protected content, the system should force a deliberate handoff from recommendation to execution.
For deeper examples of how trust boundaries fail in practice, see The 52 NHI Breaches Report and Vercel Context.ai OAuth Supply Chain Breach, which show how third-party integrations can widen exposure when permissioning is too loose.
Risk and Threat Considerations
Conversational AI becomes risky when the system treats generated text as harmless while the underlying retrieval and tool paths still have real access. The main failure mode is data overexposure, where a user receives content they were never meant to see, or where a compromised prompt or connector turns the assistant into an exfiltration channel for secrets, personal data, or internal records.
Failure mechanism: Excessive retrieval scope, weak tool authorization, or indirect prompt injection causes the model to assemble and disclose content outside the user’s legitimate boundary, especially when external connectors can read or forward data.
Impact: The result can be confidentiality loss, regulatory exposure, customer trust damage, and a much larger blast radius than a normal application bug because the assistant can aggregate and contextualise material from multiple systems at once.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Non-Human Identity Top 10, OWASP Agentic AI Top 10 and OWASP API Security Top 10 address the attack and risk surface, while NIST SP 800-53 Rev 5 and NIST Zero Trust (SP 800-207) set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP Non-Human Identity Top 10 | NHI-02 — Secret Leakage | Sensitive data leakage through assistants maps to secret and data exposure risk. |
| NHI-05 — Overprivileged NHI | Assistant tool access can exceed the user's legitimate boundary and expose data. | |
| Recommendation — Restrict retrieval and output paths that could disclose secrets or protected data. Constrain the assistant to the minimum permissions needed for the invoking user. | ||
| OWASP Agentic AI Top 10 | ASI03 — Identity & Privilege Abuse | The question is about preventing an autonomous assistant from exceeding permitted access. |
| Recommendation — Bind agent actions and data access to explicit authorization checks before execution. | ||
| NIST SP 800-53 Rev 5 | AC-6 — Least Privilege | Preventing leakage requires limiting what the assistant and its connectors can access. |
| IA-5 — Authenticator Management | Leakage often involves exposed credentials, tokens, or other sensitive identity material. | |
| Recommendation — Limit assistant and tool permissions to the minimum required for each task. Rotate and protect credentials used by assistant integrations and connectors. | ||
| NIST Zero Trust (SP 800-207) | Zero Trust Architecture | The answer depends on continuous authorization and tightly bounded trust between user, model, and tools. |
| Recommendation — Enforce verification and segmentation at every model-to-tool and model-to-data boundary. | ||
| OWASP API Security Top 10 | API5 — Broken Function Level Authorization | Model-triggered actions and tool calls can execute functions the user should not reach. |
| API1 — Broken Object Level Authorization | Retrieval and disclosure fail when the assistant can access objects outside the user's rights. | |
| Recommendation — Authorize every assistant-invoked function separately from the chat session. Check object-level access on every retrieved record before returning it. | ||
Practitioner Guidance
What to prioritise: Start with retrieval and action boundaries before tuning prompts or adding policy text. If the model can see too much, no wording change will make leakage safe.
What to verify: Test the assistant with users who should not have access to specific records, and confirm that both answers and tool outputs stay inside the same authorisation boundary as the source systems. Also verify that logging captures what was requested, what was retrieved, and what was actually disclosed.
Common mistake: Teams often harden the model prompt but leave connectors, search scope, or outbound integrations broad. That usually preserves the exfiltration path even when the wording looks restrictive.
Practitioner takeaway: Treat conversational AI as a privileged access path, not a chat feature, and make every sensitive disclosure or side effect pass through the same least-privilege and confirmation discipline you would require for any other high-impact system.
Related resources from NHI Mgmt Group
- How should security teams prevent sensitive data from leaking through AI prompts and copilots?
- How should security teams prevent LLM memory from leaking sensitive data?
- How should security teams prevent sensitive data leaks when users send email in Gmail?
- How should security teams prevent sensitive data from being copied into personal cloud and shadow AI accounts?
Deepen Your Knowledge
Reviewed and updated by the NHIMG editorial team on September 26, 2026.
NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org