Join our Newsletter — 33% off our NHI Course

What breaks when system prompts contain secrets or access rules?

The prompt stops being harmless configuration and becomes a live control surface. Once secrets, routing logic, or policy text sit inside the model, attackers can probe for them, infer behaviour, and use that knowledge to bypass safeguards or reach sensitive data.

Why system prompts stop being “just instructions” once secrets or policy are embedded

System prompts are supposed to shape behaviour, not expose protected material. When you place credentials, routing logic, tenant rules, or security policy text inside the prompt itself, you create a durable, queryable control surface. That turns the prompt into something attackers can interrogate, not merely something the model reads.

The practical distinction matters because model behaviour is not a secure vault boundary. Any content the model can condition on may be reflected, inferred, or indirectly exploited through prompt injection, extraction attempts, or behavioural probing. For teams handling sensitive instructions, OWASP Cheat Sheet Series is a useful reference for the broader control principle: do not let sensitive material become part of the application’s exposed runtime context.

That is why secret-bearing prompts tend to fail in three ways at once: confidentiality is weakened, policy becomes easier to reverse engineer, and the model can be tricked into treating hidden instructions as an attack target. The problem is not only leakage, but also the fact that hidden rules are often no longer reliably hidden once the model can be induced to talk about its own state.

What attackers can do with leaked prompt content

Once system prompts contain secrets or access rules, attackers can use them to learn how the application is wired, where trust boundaries sit, and which answers or actions are likely to succeed. That includes discovering internal endpoint names, role checks, fallback paths, tool-use logic, and language that reveals how the assistant decides what is permitted.

The second-order risk is bypass. If the prompt reveals the policy model, an attacker can craft inputs that exploit the exact structure of the rules, for example by triggering exceptions, testing contradictory instructions, or prompting the model to restate or reinterpret hidden constraints. When the prompt is carrying access logic, the attacker is effectively studying the control plane.

This is also why OWASP’s Non-Human Identity Top 10 is relevant by analogy when prompts describe machine access paths, because hidden credentials and privilege logic are no longer static text, they are part of the operating identity boundary. If those values leak, the exposure is often broader than the original prompt content suggests.

For secret handling, the safer pattern is to keep privileged material outside model-visible context and resolve it through bounded services or delegated lookups. NHIMG’s Secrets Management Guide is directly relevant here because it treats secret separation, rotation, and secretless access as a design choice, not a cleanup step after exposure.

How to design prompts so sensitive rules do not become an attack surface

The key design rule is to keep prompts descriptive, not authoritative. A prompt may summarise behaviour, but it should not contain reusable secrets, durable allowlists, or brittle policy text that an attacker can mine. Where access decisions are needed, they should be enforced by systems that can authenticate, authorise, log, and revoke independently of the model.

In practice, that means separating instruction from enforcement. The model can be told what kind of decision to request, but the actual decision should come from a control layer that can be audited and changed without rewriting the prompt. That reduces the blast radius if the model is probed or manipulated, because the attacker learns less and gains less.

For machine-access cases, the strongest pattern is to move away from long-lived prompt-embedded secrets and toward scoped, short-lived credentials or external policy checks. API Key Management Guide is useful when the exposed material is an access token or API key, while the Secret Sprawl Challenge is the better fit when the issue is hidden copies of secrets and credentials spreading into places they should never reach.

For broader architecture, NHIs should be treated as governed actors with their own lifecycle, not as strings embedded in prompts. That makes revocation, rotation, and least privilege possible even when the assistant itself is compromised or persuaded to reveal context.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

OWASP Non-Human Identity Top 10 and MITRE ATT&CK address the attack and risk surface, while NIST SP 800-53 Rev 5, OWASP ASVS and CIS Controls v8 set the governance and control requirements practitioners need to meet.

Framework Control / Reference Relevance
OWASP Non-Human Identity Top 10 NHI-02 — Secret Leakage Secrets in prompts create direct leakage exposure.
NHI-07 — Long-Lived Secrets Prompt-embedded secrets are often durable and reusable.
NHI-05 — Overprivileged NHI Prompted access rules can encode excessive authority for machine actors.
Recommendation — Remove sensitive values from prompts and keep them in bounded secret stores. Replace durable prompt secrets with short-lived, scoped credentials. Scope machine access to least privilege and revoke unused permissions.
NIST SP 800-53 Rev 5 IA-5 — Authenticator Management Prompt secrets behave like authenticators and require lifecycle control.
AC-6 — Least Privilege Access rules in prompts should not grant broader access than needed.
Recommendation — Manage secret lifecycle separately from model instructions. Enforce least privilege outside the model and limit default access.
OWASP ASVS V14 — Data Protection Sensitive values in prompts are a data protection concern.
Recommendation — Keep confidential values out of user- and model-visible context.
MITRE ATT&CK T1552 — Unsecured Credentials Embedded secrets can be discovered and reused by attackers.
Recommendation — Hunt for credential exposure and rotate any leaked secrets immediately.
CIS Controls v8 CIS-3 — Data Protection Prompt secrets are sensitive data needing protection and minimisation.
Recommendation — Classify and protect prompt content that contains sensitive information.

Practitioner Guidance

What to verify: Check whether any system prompt, template, agent instruction, or routing rule contains values that would be damaging if disclosed, copied, or replayed. If the answer is yes, move that material out of the prompt and into a control that can be rotated, audited, and independently enforced.

What to prioritise: Protect the policy boundary before tuning model output quality. A prompt that is easy to probe but hard to override is still a leak path if it exposes secrets, entitlement logic, or internal decision trees.

Common mistake: Teams often treat prompt secrecy as equivalent to access control. It is not. If the model can condition on the rule, a determined user can often infer enough of it to shape abuse, even without a literal plaintext dump.

Practitioner takeaway: If a prompt would be dangerous to disclose in an incident report, it does not belong in the prompt; keep sensitive instructions and credentials outside the model, then enforce them through an auditable control layer.