Because eval interprets attacker-influenced strings inside the host runtime, not in a safe content layer. Even when built-ins are restricted, the expression may still reach dangerous language features or runtime objects. The risk is not arithmetic itself, but letting untrusted model output become executable Python inside the application process.
Why unsafe eval turns model output into executable code
The core problem is that eval does not interpret text as a bounded answer, it asks the host runtime to execute that text as code. In an LLM integration, that collapses the boundary between untrusted model output and trusted application logic, so a prompt-influenced string can become real control flow inside the process.
That is why the risk is broader than “bad math.” Once the model can shape the expression, it may reach file operations, imports, introspection, process objects, or other runtime features that were never meant to be exposed through a content channel.
What makes the execution path dangerous in practice
Unsafe eval becomes especially risky when developers treat “restricted built-ins” as a sufficient guardrail. Runtime objects, closures, globals, helper functions, and language-specific escape hatches can still expose powerful capabilities, and the exact escape route depends on the interpreter and wrapper code.
In other words, the danger is not only that an attacker can ask for a malicious command. The danger is that the application has already handed the attacker a programmable surface inside the same security context as the service, which means compromise can extend to data access, outbound requests, secret handling, or system modification.
That makes AI infrastructure workload identity a useful companion concept here, because once code execution exists, the next question is what that process identity can reach and impersonate.
It also overlaps with LLM Provider API Key Security and LLMjacking Guide, since code execution often becomes a path to reading provider keys, calling upstream models, or abusing cloud AI credentials already available to the application.
How to think about controls when eval is involved
The right control model is to avoid letting model output enter an execution engine at all. If you need arithmetic, configuration, templates, or structured transforms, use purpose-built parsers, allowlisted operations, or a constrained expression language rather than general-purpose code evaluation.
Where execution is unavoidable, isolate it aggressively and assume the evaluated text is hostile. That means separating runtime privileges, limiting filesystem and network access, removing direct access to secrets, and treating any helper object exposed to the evaluator as part of the attack surface.
For broader governance of the surrounding AI stack, NIST AI Risk Management Framework is relevant because it frames unsafe automation as a controllable system risk rather than a coding convenience.
OWASP Agentic AI Top 10 also maps well to this problem when the eval path sits inside an agent or tool-using workflow, because runtime privilege abuse and unexpected execution are exactly the kinds of failures that turn a model mistake into a security event.
Risk and Threat Considerations
Unsafe eval creates a direct code execution path, so the main risk is not incorrect output but hostile output being executed with application privileges. In LLM integrations, that can turn a prompt injection, poisoned instruction, or malformed tool response into a server-side compromise path.
Failure mechanism: attacker-influenced text reaches eval, the runtime resolves more than the intended expression, and the evaluated code gains access to process objects, files, environment data, network calls, or imported modules.
Impact: the blast radius depends on the host process, but commonly includes secret theft, arbitrary command execution, unauthorized API calls, data exfiltration, and lateral movement through whatever credentials the application already holds.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
NIST SP 800-53 Rev 5 and OWASP ASVS set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| NIST SP 800-53 Rev 5 | IA-9 — Service Identification and Authentication | LLM integrations often run as services that can be abused after code execution. |
| AC-6 — Least Privilege | Unsafe eval becomes far worse when the process holds excess privileges. | |
| SI-10 — Information Input Validation | The issue begins when untrusted text is accepted as executable input. | |
| Recommendation — Enforce service authentication and constrain runtime credentials exposed to executed code. Minimize the process privileges available to any code path that can execute dynamic input. Validate and constrain inputs before they reach any execution-capable sink. | ||
| OWASP ASVS | V15 — Secure Coding and Architecture | Unsafe eval is a secure design flaw in application architecture. |
| Recommendation — Remove dynamic code evaluation from security-sensitive application paths. | ||
Practitioner Guidance
What to verify: confirm whether the system ever evaluates model output, prompt text, or downstream tool text with eval, exec, or equivalent dynamic execution. If yes, treat it as a security defect unless the input is fully trusted and the execution context is isolated by design.
Decision rule: if the task can be solved with parsing, templating, dispatch, or a constrained expression evaluator, choose that path instead of a general runtime evaluator. Reserve dynamic execution for cases where the code itself is the intended trusted artifact, not the model output.
Practitioner takeaway: the key judgment is to protect the execution boundary, not just the prompt boundary, because once untrusted text can execute in-process, the model is no longer generating content, it is steering privileged code.
Related resources from NHI Mgmt Group
- Why do unsafe XSLT processing settings create remote code execution risk in metadata catalog platforms?
- Why do unsafe SpEL expressions create both remote code execution and denial-of-service risk in Spring applications?
- Why do direct LLM integrations create governance risk?
- Why do direct integrations to a single LLM provider create reliability risk in enterprise AI systems?
Deepen Your Knowledge
Free weekly newsletter
Subscribe to the NHI & AI Identity Journal
The latest on NHI and Agentic AI security – articles, research, breaches, news and events every week.
Bonus 33% off our NHI Course when you subscribe.
Reviewed and updated by the NHIMG editorial team on October 8, 2026.
NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org