TL;DR: An LLM remote code execution path emerges when function-call parsing and unsafe evaluation let a prompt trigger arbitrary code on the host, even when the model itself cannot execute code, according to CyberArk analysis. The real boundary is the integration layer, where prompt handling becomes a software execution path, not an identity or policy problem.
At a glance
What this is: CyberArk examines how an LLM integration can be driven from prompt handling into arbitrary code execution when external function-call logic and unsafe eval are combined.
Why it matters: This matters to IAM and security teams because the risk sits in the trust boundary between model output, tool invocation, and host execution, which is where agentic and NHI controls often fail first.
Context
Large language models do not execute code on their own, but surrounding application logic can turn model output into code execution. In this case, the security problem is not the model alone. It is the external parser, the function-call convention, and the downstream execution path that together create a control gap around prompt handling.
For identity practitioners, the important question is how much authority an LLM integration is allowed to exercise once it can request tools or emit structured output. That is where model behaviour starts to intersect with NHI-style control boundaries, even when the underlying risk is really software execution safety.
The article is a practical example of how unsafe tool integration can collapse the distinction between text generation and privileged action. That pattern is atypical for ordinary LLM chat use, but it becomes common as developers extend models into task execution and function calling.
Key questions
Q: What breaks when LLM output is treated as trusted input?
A: When LLM output is treated as trusted input, the organisation loses the ability to separate generation from execution. Unsafe or manipulated responses can reach backend systems, disclose data, or alter business logic without a second policy decision. The failure is not the prompt itself, but the absence of a gate before action occurs.
Q: Why does unsafe eval create code execution risk in LLM integrations?
A: Because eval interprets attacker-influenced strings inside the host runtime, not in a safe content layer. Even when built-ins are restricted, the expression may still reach dangerous language features or runtime objects. The risk is not arithmetic itself, but letting untrusted model output become executable Python inside the application process.
Q: How can security teams reduce tool-misuse risk in LLM applications?
A: Use explicit authorization for every action a model can trigger, and separate content generation from state-changing operations. The model should propose, but a policy layer should decide whether the action, parameters, and target are acceptable. This is especially important when tool calls can reach files, shells, APIs, or other sensitive systems.
Q: How do you know when an LLM integration is too close to arbitrary code execution?
A: You are too close when a prompt can influence a structured instruction that the host executes with little or no review. Warning signs include direct use of eval, loose JSON parsing, and tool dispatch that trusts model-generated parameters. If the model can change system state, the integration needs stricter containment and logging.
Technical breakdown
How function-call parsing becomes an execution path
The integration described in the article uses a system prompt to tell the model which tools exist and how to format a function call. A parser then extracts a JSON block from the model output and maps that block to a local function. That design makes the model’s text output a control signal for application logic. Once the application treats model output as a trusted instruction, the model no longer needs to execute code itself. The security boundary has moved into the host application, where parsing, validation, and dispatch determine whether a benign prompt or a maliciously shaped response reaches execution.
Practical implication: Treat model output as untrusted input and validate every function call before dispatching it.
Why unsafe eval turns a utility function into RCE
The calculate function in the article uses Python eval with a restricted namespace and no built-ins. That looks controlled, but eval remains dangerous because Python’s object model and import mechanics can often be abused to regain code execution. A sandbox that blocks obvious built-ins can still fail if the evaluated expression can reach powerful runtime objects or modules indirectly. In other words, the vulnerability is not the arithmetic feature. It is the decision to interpret attacker-influenced strings as executable Python in-process, where the application inherits the full risk of the interpreter.
Practical implication: Keep untrusted expressions out of eval and isolate any code execution in a hardened boundary.
How LLM integrations create tool-misuse risk
The article shows that the attack succeeds when the model can be coaxed into emitting the exact JSON structure the application expects for a function call. That is a classic tool-misuse pattern: the attacker does not need the model to understand the consequence, only to produce a format the host trusts. This is not the same as prompt injection alone. It is prompt manipulation plus deterministic downstream execution. As integrations become more capable, the control problem shifts from content moderation to authorisation of actions, parameters, and execution context.
Practical implication: Gate tool use with explicit policy checks on action, arguments, and execution context, not just prompt content.
Threat narrative
Attacker objective: The attacker wants to convert model-mediated input into code execution on the host system.
- Entry occurs when an attacker supplies a prompt that steers the model toward emitting a crafted function-call payload.
- Credential access is not the core stage here; instead, the hostile payload reaches the host parser as trusted structured output.
- Escalation happens when the application passes attacker-shaped parameters into eval and the Python runtime interprets them as executable code.
- Impact is arbitrary command execution on the server, which can then be used for file creation, post-exploitation, or broader system compromise.
Breaches seen in the wild
- Moltbook AI agent keys breach: Moltbook breach exposed 1.5M AI agent keys.
- McKinsey AI platform breach: McKinsey AI platform hack exposed 46M chats and sensitive data.
Read and download The State of NHI & AI Agent Breach Report 2026, covering 200+ breaches impacting Non-Human Identities including AI Agents.
NHI Mgmt Group analysis
Prompt handling is not the security control. Once an LLM integration turns text into a function invocation, the control boundary shifts from natural language to executable application logic. That boundary is only as strong as the parser, validator, and dispatcher that sit between the model and the host system. The practitioner implication is that tool invocation must be governed like code execution, not like chat content moderation.
Unsafe eval creates an execution bridge, not just a coding mistake. The article’s core failure is that a convenience function becomes a local interpreter for attacker-influenced input. That pattern matters because many teams still treat model outputs as soft signals rather than high-risk execution inputs. The implication is that runtime authorisation must be explicit anywhere model text can change host behaviour.
Tool misuse is the sharper risk than prompt misuse. A model can be manipulated into producing a valid function payload even when it has no malicious intent of its own. That means security teams should measure whether the surrounding system can constrain what a model is allowed to ask the host to do, not whether the prompt looked suspicious. The practitioner implication is to design for hostile outputs, not friendly intent.
LLM integrations expose a trust boundary between language and action. The article shows that the dangerous point is where natural language becomes structured instruction and then becomes system effect. This is where OWASP-AGENTIC style concerns around tool misuse intersect with classic software execution risk. The practitioner implication is to separate model generation from privileged action with enforceable policy checks and execution containment.
Prompt-to-code translation is now a named security pattern. It describes systems where structured model output is directly consumable by downstream code paths, creating a reusable attack pattern across LLM orchestration stacks. That pattern is broader than a single product or model family and will recur wherever developers let language models trigger tools. The practitioner implication is to treat every model-tool bridge as a potential execution surface.
From our research library:
- AI-related credential leaks surged 81.5% year-over-year in 2025, with the surrounding AI infrastructure leaking 5x faster than core LLM providers, according to the State of Secrets Sprawl 2026.
- Only 44% of developers are reported to follow security best practices for secrets management, exposing a significant developer behaviour gap, according to the State of Secrets in AppSec.
- Read next: AI Agent Observability, Audit and Incident Response Guide
What this signals
Prompt-to-code translation: systems that convert model output directly into function calls create a repeatable execution surface, not just a content-safety issue. The operational signal is simple: if a model can trigger code paths, then the security model has moved from conversation control to action control.
The next governance step is to separate model intent from host authority. That means explicit policy checks, strong schema validation, and detailed auditability around every tool invocation so that a model cannot turn structured output into privileged execution.
For practitioners
- Constrain function-call parsing Reject any model output that does not match a narrowly defined schema and approved action list before it reaches execution logic.
- Remove eval from host workflows Replace in-process eval with non-executable calculation paths or isolated services that cannot reach ambient interpreter capabilities.
- Separate model output from privilege Require an explicit policy decision before any tool invocation that changes state, accesses files, or reaches external systems.
- Instrument tool-use logging Record the originating prompt, emitted function call, argument values, and resulting host action so abuse can be traced after execution.
- Test for prompt-to-code abuse Red-team the integration with benign-looking prompts that try to elicit structured calls, unsafe arguments, or chained tool actions.
Key takeaways
- This case shows that LLM risk escalates sharply when model output is allowed to drive executable logic in the host application.
- The failure is not the model alone, but the integration pattern that trusts structured output and feeds it into unsafe code paths.
- Containing the risk requires separating generation from execution, removing eval, and validating every tool call before it is dispatched.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Agentic AI Top 10, OWASP Non-Human Identity Top 10 and MITRE ATT&CK address the attack and risk surface, while NIST AI RMF sets the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP Agentic AI Top 10 | ASI02 — Tool Misuse | The article centers on an LLM being manipulated into unsafe tool execution. |
| ASI05 — Unexpected Code Execution | Unsafe eval turns model-triggered input into host code execution. | |
| Recommendation — Restrict agent tool use to approved actions and validate every argument before execution. Eliminate in-process code evaluation for model-influenced input and isolate execution paths. | ||
| OWASP Non-Human Identity Top 10 | NHI-04 — Insecure Authentication | Model-to-tool trust is weakly authenticated and can be abused as a control boundary. |
| Recommendation — Authenticate and authorise tool requests separately from model output before any state change. | ||
| MITRE ATT&CK | TA0002;TA0004;TA0006 — Execution; Privilege Escalation; Credential Access | The attack path moves from malicious input to code execution and then broader host abuse. |
| Recommendation — Map the integration to execution and privilege-escalation paths, then harden the host against post-exploitation. | ||
| NIST AI RMF | GOVERN — AI Governance and Accountability | The issue is governance of AI-triggered actions and accountability for host-side effects. |
| Recommendation — Define accountability and approval boundaries for every AI-driven action that can change system state. | ||
Key terms
- Function-call injection: A pattern where an attacker manipulates model output so it is accepted as a valid tool request by the host application. The risk is not just prompt wording. It is the trust placed in structured output that can trigger privileged actions or unsafe execution.
- Unsafe eval: Use of an interpreter function that executes dynamically supplied expressions inside the application process. In LLM systems, this becomes dangerous when model-influenced text reaches eval or similar mechanisms, because the application is then responsible for containing code it never truly controlled.
- Tool Misuse: Tool misuse occurs when an agent uses an allowed integration in a way that exceeds its intended task, scope, or risk tolerance. The problem is often not access alone but the combination of valid credentials, broad permissions, and unbounded action sequencing.
- Prompt-to-code translation: A design pattern in which model-generated text is directly converted into executable or state-changing instructions. It is a high-risk boundary because the semantic gap between language and action disappears, making validation, authorisation, and containment the primary security controls.
Deepen your knowledge
NHI governance, agentic AI identity, and machine identity lifecycle are core topics in our NHI Foundation Level course, the industry's only accredited NHI security programme. If you are building or maturing an IAM programme, it is worth exploring.
Published by the NHIMG editorial team on May 29, 2026.
Updated on October 8, 2026.
NHI Mgmt Group, the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org