Because prompt visibility does not control runtime behaviour. A system prompt can look safe while the assistant still discloses data through hidden triggers, formatted output, or client-side rendering. The practical risk is not only what the prompt says, but what the deployed assistant can do with user input during execution.
Why prompt visibility does not eliminate leakage
Prompt-visible assistants can still leak data because the prompt is only one layer of control, while the deployed assistant has runtime behavior. Even a well-written system prompt does not stop hidden instructions, prompt injection, unsafe tool calls, over-broad retrieval, or output paths that expose content after generation. The security question is whether the assistant can be made to reveal data, not whether the prompt looks safe on inspection.
That gap matters most when the assistant can see user content, retrieve documents, or transform information into a new format. A visible prompt may reassure reviewers, but it does not govern every branch of execution. If the assistant can read sensitive context, a malicious input can still shape what is recalled, summarized, rendered, or copied into a response.
For teams building or reviewing assistants, the right mental model is not “the prompt is visible, therefore the system is safe.” It is “the assistant must be constrained at retrieval, generation, and rendering.” Prompt text helps set policy, but leakage risk lives in the execution path.
Where leakage actually happens in deployed assistants
The most common failure point is not the instruction text itself, but the paths that surround it. Retrieval systems can surface material the user should not see, especially when permissions are missing or weakly enforced. Output can also leak through quoting, summarization, or structured formatting that bypasses normal human review. In some deployments, the client layer adds another exposure point if rendered content is copied into logs, DOM elements, exports, or downstream integrations.
That is why permissioning and retrieval design matter as much as prompt engineering. A Permission-Aware RAG Guide is directly relevant here because it addresses the core issue: sensitive content must be filtered before it reaches the model, not merely discouraged in prompt wording. If retrieval is over-broad, the assistant can expose data even when the prompt is conservative.
Visible prompts also do not neutralize indirect attacks. A crafted message can still manipulate context, and a benign-looking output can still contain copied secrets, internal identifiers, or confidential snippets. That is why leakage prevention needs controls around data access, not just instructions about behavior.
Why reviewable prompts can still be dangerous at scale
Reviewability is useful, but it can create false confidence when organizations treat prompt inspection as a proxy for system safety. The larger the deployment, the more likely the assistant will sit in front of diverse inputs, diverse permissions, and diverse downstream consumers. At that point, the practical question becomes blast radius: which data can the assistant touch, which users can trigger it, and where does its output flow next?
Prompt visibility does not answer those questions. A system can be transparent and still unsafe if it combines broad retrieval, weak authorization, and permissive output handling. In other words, the prompt may be visible while the data path remains invisible.
The lesson from real deployments is straightforward: disclosure risk often comes from ordinary product behavior, not exotic model failures. The EchoLeak (Microsoft 365 Copilot) 2025 example shows how a crafted input can drive unintended disclosure from assistant context without the user doing anything unusual. That is the sort of failure that prompt review alone will not catch.
Risk and Threat Considerations
Prompt-visible assistants increase the risk of overtrust, because reviewers may assume that readable instructions equal controllable behavior. The real exposure is that sensitive context can be pulled into the model, manipulated by hostile input, and released through ordinary output channels, even when the prompt itself appears disciplined.
Failure mechanism: The assistant processes user input and hidden context at runtime, so injected instructions, broad retrieval, or unsafe rendering can cause disclosure despite a safe-looking system prompt.
Impact: Sensitive text, internal documents, credentials, or customer data can leave the intended trust boundary through chat responses, citations, formatted output, logs, or client-side presentation.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Non-Human Identity Top 10, OWASP Agentic AI Top 10 and OWASP API Security Top 10 address the attack and risk surface, while NIST SP 800-53 Rev 5 sets the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP Non-Human Identity Top 10 | NHI-02 — Secret Leakage | Prompt-visible assistants can still disclose sensitive data through runtime paths. |
| NHI-05 — Overprivileged NHI | Leakage often follows excessive access to data and tools at execution time. | |
| Recommendation — Block sensitive retrieval and redact secrets before the assistant can emit them. Reduce the assistant’s data and tool access to the minimum required. | ||
| OWASP Agentic AI Top 10 | ASI02 — Tool Misuse | Runtime behavior can expose data when tools or actions are invoked unsafely. |
| Recommendation — Constrain tool execution so the assistant cannot retrieve or export unauthorized data. | ||
| OWASP API Security Top 10 | API1 — Broken Object Level Authorization | Assistant retrieval and downstream data access fail when object-level permissions are not enforced. |
| Recommendation — Enforce object-level authorization on every data object the assistant can reach. | ||
| NIST SP 800-53 Rev 5 | AC-6 — Least Privilege | Prompt visibility does not matter if the assistant is over-privileged at runtime. |
| Recommendation — Limit the assistant to the minimum permissions needed for its task. | ||
Practitioner Guidance
What to verify: Check where the assistant can read from, not just what it is told to do. If retrieval, tool access, or rendering can reach sensitive data, the prompt is not the primary control boundary.
Decision rule: If a prompt-visible assistant can access data it should not disclose, fix authorization, retrieval filtering, and output handling before relying on prompt reviews. Treat prompt inspection as a support control, not the safety guarantee.
What practitioners underestimate: Visibility often improves auditability but not containment. The most important question is whether the deployed assistant can be induced to move protected data across a boundary during execution.
Practitioner takeaway: A readable prompt can help you understand intent, but only runtime controls determine whether the assistant can leak data in practice.
Related resources from NHI Mgmt Group
- Why do approved AI tools still create data leakage risk?
- Why do valid API calls still create data leakage risk?
- Why does email still create so much data leakage risk in organisations with mature security controls?
- Why do AI assistants create data leakage risk when file permissions are updated in enterprise environments?