TL;DR: LLM abuse now spans prompt injection, prompt leaking, and jailbreaking, while offensive-capable models lower the barrier to generating malicious code and attack guidance, according to CyberArk. The governance gap is no longer theoretical: identity security and layered controls now matter as much as model safety.
At a glance
What this is: This is CyberArk’s analysis of how unguarded LLMs expand attack paths through prompt injection, prompt leaking, jailbreaking, and openly accessible offensive-capable models.
Why it matters: It matters because security teams now have to govern who and what can access AI systems, how outputs are constrained, and where identity controls sit in the defensive stack.
Context
Large language models can be manipulated through prompts that change output behaviour, bypass safeguards, or elicit prohibited content. The article treats that as an identity and access problem as much as a model-safety problem, because the same systems are being used to request harmful instructions, code, and operational guidance.
The governance gap is that many teams still think of AI risk as content moderation alone. For IAM and NHI programmes, the real question is how model access, login pathways, and output permissions are governed when the system can be used by both legitimate users and offensive actors.
Key questions
Q: What breaks when LLMs are exposed to unrestricted prompt access?
A: Unrestricted prompt access turns the model into a writable execution surface, so users can steer outputs toward unsafe instructions, hidden data, or unauthorized actions. The failure is not only bad content. It is a missing boundary between who may ask, what they may ask for, and what the system is allowed to return.
Q: Why do offensive-capable AI services increase abuse risk?
A: They lower the cost of attack preparation by packaging harmful guidance, code, and exploit logic behind normal access flows. That means abuse no longer depends only on advanced technical skill. It depends on who can reach the service and whether the provider has limited the kinds of outputs the model can generate.
Q: What are the signs that guardrails are failing in an LLM application?
A: Guardrails are likely failing when the model returns unsafe, off topic, or unstructured responses that bypass policy checks. Other warning signs include inconsistent formatting, missing validation of sensitive outputs, and corrective actions that do not trigger when criteria are not met. If users can still provoke harmful or noncompliant replies, the control is too weak or misconfigured.
Q: How should security teams govern generative AI tools that connect to core systems?
A: Treat them as non-human identities with lifecycle, access, and telemetry requirements. Assign an owner, limit privileges to the exact task, log every data flow they can trigger, and revoke access immediately when the business need ends. If a tool cannot be inventoried or monitored, it should not be connected to sensitive systems.
Technical breakdown
Prompt injection, prompt leaking, and jailbreaking
Prompt injection relies on crafted inputs that look legitimate enough to steer model behaviour toward disallowed outputs, unauthorized actions, or data exposure. Prompt leaking uses training examples, embedded instructions, or model memory to infer hidden context and bypass intended restrictions. Jailbreaking is the deliberate attempt to override system safeguards and developer instructions through prompt structure alone. These are different techniques, but they share one property: the model responds to language as an execution surface, not just as text. That makes input handling, output policy, and instruction hierarchy part of the security boundary.
Practical implication: treat prompts as untrusted input and restrict what the model can disclose, generate, or execute.
Offensive-capable LLMs change the attack economy
The article’s WhiteRabbitNeo example shows a model designed for offensive research, with login-based access and no general-purpose safeguards. That changes the economics of abuse because attackers do not always need to defeat guardrails if the service already returns malicious code, exploit logic, or attack instructions by design. In other words, the threat is not only model compromise but also model availability. Once offensive guidance is packaged as a service, lower-skill actors can access techniques that previously required tooling, expertise, or more direct compromise.
Practical implication: assess whether AI services are reducing attacker effort by packaging harmful capability behind ordinary authentication flows.
Identity controls sit in front of model risk
The article’s strongest operational point is that AI safety cannot be separated from identity governance. Access to an AI service, the legitimacy of the calling identity, and the scope of permitted use all shape whether the system becomes a helpful assistant or a security problem. That is especially true where login providers, API access, or shared accounts make attribution weak. Model safeguards help, but they do not replace entitlement checks, session visibility, or usage governance. The security boundary is wider than the model prompt alone.
Practical implication: connect model access to identity, entitlement, and audit controls rather than treating AI risk as a standalone content problem.
Threat narrative
Attacker objective: The attacker wants scalable, low-friction access to harmful instructions and code that can be reused for fraud, intrusion, or malware development.
- Entry occurs when an attacker uses ordinary prompt access or a login-backed offensive AI service to reach model capability without having to exploit a traditional software vulnerability.
- Credential or access abuse follows when the same authenticated path is used to request malicious code, bypass safeguards, or retrieve prohibited operational guidance.
- Impact is the generation of usable attack content, including malware logic, exploit steps, and instructions that reduce the skill barrier for abuse.
Breaches seen in the wild
- Samsung ChatGPT leak 2023: Samsung staff pasted chip source code and meeting notes into ChatGPT weeks after it was allowed, leading Samsung to restrict generative AI tools.
- McKinsey AI platform breach: McKinsey AI platform hack exposed 46M chats and sensitive data.
Read and download The State of NHI & AI Agent Breach Report 2026, covering 200+ breaches impacting Non-Human Identities including AI Agents.
NHI Mgmt Group analysis
LLM abuse is now an identity problem, not only a content-safety problem. The article shows that prompt access, login access, and output control all shape whether a model becomes an exposure point. Guardrails matter, but they sit downstream of entitlement and session governance. Practitioners should treat AI access as part of the identity control plane, not as a separate policy conversation.
Prompt injection works because the model accepts language as an execution path. The security failure is not simply that bad prompts exist, but that systems often do not distinguish trusted instructions from user-supplied instructions with enough rigor. That collapses the boundary between request and action. Practitioners should assume that model interaction surfaces need stronger authorization logic than ordinary application text boxes.
Offensive-capable models create a new distribution channel for attacker tradecraft. When a service can generate code, attack steps, or exploitation guidance on demand, the barrier shifts from skill to access. That expands the population that can participate in abuse, which is why model availability now deserves the same attention as model safety. Practitioners should evaluate AI services by the harm they can operationalise, not by the marketing language around research or testing.
Identity security becomes the control that makes AI governable. The article’s closing logic is correct: layered defence starts with authenticity, integrity, and access discipline. Without those, prompt filters only trim symptoms. The field should stop discussing AI as though model safeguards alone define the boundary, because governance failures emerge the moment access, misuse, and attribution are uncontrolled. Practitioners should anchor AI programmes in identity-led controls.
AI without guardrails is a named governance concept for this category. It describes the point at which model capability, permissive access, and weak behavioural limits combine to make harmful output easy to produce. That is the condition teams must design against across human users, service accounts, and AI-enabled workflows. Practitioners should build policies around constrained access, not assumed intent.
From our research library:
- AI-related credential leaks surged 81.5% year-over-year in 2025, with the surrounding AI infrastructure leaking 5x faster than core LLM providers, according to the State of Secrets Sprawl 2026.
- Read next: AI Agent Authorisation Guide
What this signals
AI attack governance now needs the same discipline applied to high-risk NHI access. If a model can generate malicious code or unsafe instructions, the access question is no longer optional. Organisations should expect AI services to be reviewed like any other privileged capability, with entitlement scope, session visibility, and auditability built into the operating model.
Prompt surfaces have become a practical trust boundary. The model is not the only issue. The request path, the identity behind the request, and the downstream use of the output all determine whether an AI system becomes an asset or an attack amplifier.
For practitioners
- Govern AI access as an identity-controlled service Tie LLM access to named identities, approved use cases, and auditable sessions instead of letting model access float outside IAM and PAM oversight.
- Restrict offensive-capability exposure Separate general-purpose assistants from services intended for red-team or adversarial research, and require explicit business justification for access to offensive-capable models.
- Monitor prompt and output abuse patterns Log high-risk prompt categories, repeated jailbreak attempts, and requests for malware, exploit steps, or sensitive data extraction so reviewers can spot misuse early.
- Harden employee guidance on sensitive input Publish clear rules for what staff may never paste into an LLM, including source code, credentials, incident details, and regulated data.
- Add AI use to access review scope Include AI services, connected accounts, and shared login pathways in entitlement reviews so overbroad model access does not persist unnoticed.
Key takeaways
- Ungoverned LLM access can turn ordinary prompts into a route for generating malicious code, attack guidance, and prohibited outputs.
- The article shows how offensive-capable models can lower the barrier to abuse by packaging harmful capability behind routine authentication and normalised access.
- Security teams need identity-led controls, logging, and usage restrictions around AI services because model guardrails alone do not create a complete control boundary.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Agentic AI Top 10 and OWASP Non-Human Identity Top 10 address the attack and risk surface, while NIST AI RMF and NIST CSF 2.0 set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP Agentic AI Top 10 | ASI02 — Tool Misuse | The article focuses on AI systems being used to generate harmful outputs and attack guidance. |
| ASI09 — Human-Agent Trust Exploitation | Prompt abuse leverages user trust in the model interface to elicit prohibited behaviour. | |
| Recommendation — Constrain agent and model outputs so prompts cannot be used to generate harmful operational guidance. Reduce trust exploitation by validating user intent and limiting what the model can reveal or execute. | ||
| OWASP Non-Human Identity Top 10 | NHI-10 — Human Use of NHI | The article shows humans using AI systems as identity-bound services to obtain unsafe outputs. |
| Recommendation — Review whether people are using AI services as shortcuts to bypass normal security and approval boundaries. | ||
| NIST AI RMF | GOVERN — AI Governance and Accountability | The article argues that AI risk must be managed through governance, accountability, and oversight. |
| Recommendation — Establish accountable ownership for AI service access, acceptable use, and misuse response. | ||
| NIST CSF 2.0 | PR.AA-05 — Access Permissions, Entitlements and Authorizations | The core control issue is who can reach AI services and what they are authorised to do. |
| Recommendation — Apply entitlement reviews to AI services so access is limited to approved identities and use cases. | ||
Key terms
- Prompt Injection (Agentic): An attack where malicious instructions are embedded in content that an AI agent reads, causing the agent to execute unintended actions using its own legitimate credentials. A primary vector for agent goal hijacking and identity abuse.
- Jailbreaking: Jailbreaking is the practice of crafting prompts that persuade an AI model to ignore its safeguards and produce restricted outputs. It shows that authentication to the service does not guarantee safe behavior, which is why governance must extend beyond the chat interface.
- Offensive-Capable Model: An offensive-capable model is an AI service that can generate malicious code, exploit guidance, or other harmful content on demand. The security issue is not only what it can do, but who can access it, how that access is governed, and whether the outputs are auditable.
- AI Service Access Governance: AI service access governance is the set of identity, entitlement, and audit controls that determine who may use an AI system and for what purpose. For LLMs, it includes approved identities, session visibility, usage restrictions, and review of high-risk request patterns.
Deepen your knowledge
NHI governance, agentic AI identity, and machine identity lifecycle are core topics in our NHI Foundation Level course, the industry's only accredited NHI security programme. If you are building or maturing an IAM programme, it is worth exploring.
Published by the NHIMG editorial team on May 29, 2026.
Updated on October 8, 2026.
NHI Mgmt Group, the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org