TL;DR: OWASP’s 2025 Top 10 for LLM Applications adds new categories for excessive agency, system prompt leakage, vector weaknesses and unbounded consumption, while reworking earlier risks around prompt injection, disclosure and supply chain exposure. The update shows that AI security now hinges on identity, access and control boundaries rather than model quality alone, according to Aembit. Access review assumptions break when nonhuman actors can act, leak and chain decisions inside a single session.
At a glance
What this is: This is Aembit’s analysis of the 2025 OWASP Top 10 for LLM Applications, highlighting that the most material risks now sit at the intersection of LLM behaviour, identity, access control and untrusted inputs.
Why it matters: It matters because IAM, PAM and NHI programmes now have to govern agent permissions, prompt boundaries and runtime authority, not just model selection or data quality.
Context
LLM security is no longer just a model-quality problem. The article argues that once applications can read untrusted content, call tools and act on behalf of users, the main security boundary becomes identity, privilege and instruction separation.
The 2025 OWASP Top 10 for LLM Applications reflects that shift by reworking existing risks and adding categories for agency, leakage, vector weakness and resource abuse. For IAM teams, the practical issue is whether the controls around nonhuman actors still assume a passive system that only answers questions.
Key questions
Q: What breaks when LLMs can act with excessive agency?
A: The control model breaks when an AI agent can reach more tools, more data or more actions than the task requires. At that point, the risk is no longer limited to bad answers. The agent can send messages, change records or trigger workflows, so privilege scope and approval boundaries become the real security control.
Q: Why do prompt injection attacks create governance risk for AI agents?
A: Prompt injection creates governance risk because the model often sits in the control path between text input and tool execution. If attackers can change what the model treats as authoritative, they can influence access decisions, data exposure, or downstream actions without compromising a traditional account. That makes prompt provenance and instruction hierarchy part of AI identity governance.
Q: How should teams govern system prompts in LLM applications?
A: Teams should treat system prompts as configuration, not as a secrecy boundary. Prompts can leak, be inferred or be manipulated, so any critical control that depends on them is fragile. Keep privilege separation, authorization and sensitive business logic in external deterministic systems instead.
Q: How do organisations know if an LLM deployment is overstepping its authority?
A: Look for models or agents that can reach unrelated tools, perform high-impact actions without approval, or expose internal rules and secrets through normal interaction. Those are signs that authority is too broad and the security boundary is being enforced by the model instead of by policy.
Technical breakdown
Prompt injection exploits the instruction-data boundary
Prompt injection works because many LLM systems process instructions and data in the same channel. That means malicious text inside a user prompt, a webpage, a document or even an image can be treated as control input rather than content. Direct injection is explicit, while indirect injection hides the malicious instruction in retrieved or ingested material. Multimodal systems widen the attack surface because the model may interpret nontext inputs alongside trusted text. The core failure is not simply bad content, but the lack of a hard boundary between what the model should read and what it should obey.
Practical implication: separate instruction sources from untrusted content and enforce boundary controls outside the model.
Excessive agency turns AI agents into overprivileged actors
Excessive agency appears when an agent can reach too many tools, hold broader permissions than the task needs, or execute consequential actions without human approval. OWASP breaks this into excessive functionality, excessive permissions and excessive autonomy. In identity terms, the risk is not the model itself but the authority wrapped around it. If the agent can email, query databases or trigger workflows, its permissions become part of the attack surface. A deterministic control plane must decide what the agent can do, while the model should only operate within those externally enforced limits.
Practical implication: scope agent permissions to each task and keep authorization external to the model.
System prompt leakage reveals hidden control logic
System prompts often contain internal rules, filtering logic, permission structures and operational context. OWASP’s point is that these prompts are not a safe place to hide security controls because attackers can extract them through interaction. Once exposed, they help attackers target bypasses, craft better prompt injections and infer where privilege boundaries are weak. The important architectural lesson is that prompt text is not a reliable security layer. Any control that depends on secrecy inside the prompt is a soft control, not a durable one, and should be treated that way in the security model.
Practical implication: keep privilege separation and authorization in deterministic systems, not in prompt text.
Threat narrative
Attacker objective: The attacker’s objective is to convert LLM access into unauthorized disclosure, unwanted action or operational disruption by abusing the model’s authority.
- Entry begins when an attacker places malicious instructions into a prompt, document, webpage or embedded content that an LLM will later process. The model treats the injected instruction as input it should obey.
- Credential or authority abuse follows when the LLM has access to tools, APIs, search or downstream systems and uses those permissions to act on the attacker’s behalf. In agentic designs, the model can also leak system prompts or internal control logic that make follow-on abuse easier.
- Escalation occurs when excessive agency lets the model chain decisions, reach broader permissions than the task needs or execute high-impact actions without approval. The attack can then move from manipulation to harmful action.
- Impact is achieved when the model sends messages, modifies records, exposes sensitive information or consumes resources in ways the organisation did not intend.
Breaches seen in the wild
- McKinsey AI platform breach: McKinsey AI platform hack exposed 46M chats and sensitive data.
- Moltbook AI agent keys breach: Moltbook breach exposed 1.5M AI agent keys.
Read our 52 NHI Breaches Analysis report for a comprehensive view of breaches impacting Non-Human Identities including AI Agents.
NHI Mgmt Group analysis
Identity, not model quality, is now the primary security boundary for LLM applications. The article makes clear that the most damaging LLM risks emerge when models can read, retrieve and act across trust boundaries. That changes the governance question from how accurate the model is to who or what is allowed to influence or execute actions through it. The practical conclusion is that IAM, PAM and workload identity controls now sit inside the LLM security model, not beside it.
Access review logic fails when nonhuman actors can decide and act within the same session. Traditional governance assumes privilege is stable long enough to be observed, certified and revoked on a review cadence. Agentic systems break that assumption by requesting, using and discarding authority inside a single task. That is why review-based control alone cannot describe the risk. Practitioners need to rethink authority at issuance time, where the actual decision path begins.
System prompts are control surfaces, not secrets. The article correctly treats system prompt leakage as an architectural problem, not just an information exposure issue. When teams place security logic inside prompt text, they create a brittle control that can be surfaced, copied or manipulated by an adversary. The governance implication is straightforward: durable privilege separation belongs in deterministic systems, while prompts should be treated as mutable instructions, not trusted policy.
Ephemeral model behaviour creates an identity blast radius that conventional application security does not measure. An LLM or agent can amplify a small mistake into broad access, disclosure or action across several tools in one workflow. That makes the real unit of risk the reachable blast radius of the nonhuman actor, not the model endpoint itself. Security teams should evaluate how far a model can move across systems before human oversight can intervene.
From our research library:
- AI-related credential leaks surged 81.5% year-over-year in 2025, with the surrounding AI infrastructure leaking 5x faster than core LLM providers, according to the State of Secrets Sprawl 2026.
- Read next: Agentic AI Security Guide
What this signals
Ephemeral credential trust debt: when an LLM or agent can obtain and use authority faster than a review cycle can observe it, governance shifts from certification to issuance-time control. The security programme has to measure the reachable blast radius of nonhuman actors, not just whether the model is accurate.
The update also reinforces that prompt and tool boundaries belong in the surrounding control plane. If the model can see, infer or execute beyond its task scope, the problem is not only misuse, it is a governance design that granted the model too much operational leverage in the first place.
For practitioners
- Map nonhuman authority boundaries Inventory which tools, data stores and outbound actions each LLM or agent can reach, then remove anything not required for the specific task.
- Move authorization outside prompts Keep access decisions in deterministic policy systems and treat prompt text as untrusted instruction content, not as a control plane.
- Restrict agent permissions by task Assign the minimum permissions needed for each workflow, and separate read, write and act capabilities so a single model session cannot pivot broadly.
- Treat model outputs as untrusted Validate every LLM output before it reaches a browser, database, ticketing system or API call, especially when the model can chain steps.
- Review hidden control assumptions Check whether any security rule, filter or approval step lives only in a system prompt and replace it with enforceable controls outside the model.
Key takeaways
- LLM application risk now sits at the intersection of model behaviour, identity boundaries and tool authority, not just prompt quality or data hygiene.
- The 2025 OWASP update reflects real-world pressure from agency, leakage, vector weakness and resource abuse across the LLM lifecycle.
- Practitioners need controls that separate instructions from data, constrain agent permissions and keep authorization outside the prompt layer.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Agentic AI Top 10, OWASP Non-Human Identity Top 10 and MITRE ATT&CK address the attack and risk surface, while NIST AI RMF and NIST Zero Trust (SP 800-207) set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP Agentic AI Top 10 | ASI01 — Agent Goal Hijack | Prompt injection and agency risks map directly to agent goal hijack in LLM systems. |
| ASI02 — Tool Misuse | The article centres on models calling tools and misusing connected actions. | |
| ASI03 — Identity & Privilege Abuse | Excessive agency is fundamentally an identity and privilege problem for agents. | |
| Recommendation — Constrain agent goals and isolate instruction sources so hostile input cannot redirect execution. Restrict tool access by task and enforce approval for high-impact actions. Bind agent authority to least privilege and keep authorization external to the model. | ||
| OWASP Non-Human Identity Top 10 | NHI-04 — Insecure Authentication | AI agents and LLM applications rely on machine authentication to reach tools and data. |
| NHI-05 — Overprivileged NHI | The article repeatedly warns against broad permissions for agents and connected components. | |
| Recommendation — Use strong machine authentication and avoid shared credentials across LLM workflows. Reduce nonhuman privileges to the minimum task scope and review every high-impact entitlement. | ||
| MITRE ATT&CK | TA0006;TA0008 — Credential Access; Lateral Movement | Prompt injection and tool abuse can lead to credential exposure and movement across connected systems. |
| Recommendation — Map LLM abuse paths to credential access and lateral movement to prioritise detection and containment. | ||
| NIST AI RMF | GOVERN — AI Governance and Accountability | The article is about governing AI behaviour, permissions and accountability across the lifecycle. |
| Recommendation — Define governance for AI authority, oversight and escalation before deploying agentic workflows. | ||
| NIST Zero Trust (SP 800-207) | 3.2 — Continuous Verification | The article’s core issue is continuous trust when nonhuman actors act through multiple systems. |
| Recommendation — Apply continuous verification to nonhuman actions rather than trusting session start authentication. | ||
Key terms
- Prompt Injection (Agentic): An attack where malicious instructions are embedded in content that an AI agent reads, causing the agent to execute unintended actions using its own legitimate credentials. A primary vector for agent goal hijacking and identity abuse.
- Excessive agency: A condition where an AI system is given more operational authority than its task requires. The risk is not just poor output. It is that mistakes, manipulation, or compromise can produce destructive actions at machine speed across the systems the agent can reach.
- System Prompt Leakage: System prompt leakage is the exposure of hidden prompt content to users or attackers. The real security problem is usually not the prompt itself, but the secrets, policy logic, and internal architecture details placed inside it. If those details are sensitive, they should live in code or secrets management instead.
- Untrusted Output Handling: Untrusted output handling is the failure to validate LLM-generated content before another system acts on it. This matters because model output can contain executable commands, malicious text or incorrect instructions, and downstream systems may treat it as authoritative if no control layer intervenes.
Deepen your knowledge
NHI governance, agentic AI identity, and machine identity lifecycle are core topics in our NHI Foundation Level course, the industry's only accredited NHI security programme. If you are building or maturing an IAM programme, it is worth exploring.
Published by the NHIMG editorial team on June 6, 2026.
Updated on October 6, 2026.
NHI Mgmt Group, the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org