Because the failure point is often the interaction layer, not the codebase. Prompt injection, manipulated context, unsafe outputs, and delegated actions can all produce harm without a classic exploit. The risk comes from runtime trust decisions that code scanning and static controls do not fully model.
Why secure code can still be unsafe at runtime
Secure code reduces classic software flaws, but LLM systems fail in a different place: the live interaction between user input, retrieved context, model behavior, and downstream actions. A model can be tricked into following malicious instructions, exposing data, or producing harmful output even when the application code has no obvious vulnerability. The security boundary shifts from source code to runtime trust.
That is why code scanning alone gives a false sense of safety. A well-tested application can still accept poisoned prompts, unsafe tool calls, or manipulated context that changes what the system decides to reveal or do. The practical question is not only whether the code is correct, but whether the runtime policy around the model is strong enough to resist adversarial input and misuse.
When you evaluate LLM risk, treat the model interaction layer as an attack surface in its own right. A prompt, a memory store, a retrieval result, or an agent action can all become the point where security fails, even if the underlying application stack is clean.
Where the real failure point sits
LLMs introduce a control problem, not just a software defect problem. The model may accept instructions that were never intended by the developer, especially when untrusted text is mixed with trusted instructions or when the system delegates work to tools. That makes prompt injection, context poisoning, and unsafe delegation material risks rather than edge cases.
Runtime trust is especially fragile because the system often cannot reliably distinguish user intent, embedded instructions, and retrieved content. If the model can read a document, call a tool, or act on behalf of a user, an attacker only needs to steer those decisions once. The application may still be “secure” in the conventional sense, but the outcome can still be wrong, leaky, or destructive.
This is also why LLM failures often look like abuse rather than exploitation. The system may not crash, and there may be no code exploit to patch. Instead, the model follows a path that the designer did not intend, which is why runtime policy, input separation, and action boundaries matter as much as secure development.
Why this matters for practitioners
The security objective is to bound what the model can see, infer, and do. That means separating trusted instructions from untrusted content, constraining tool access, and deciding which outputs are allowed to trigger side effects. If you allow a model to draft, retrieve, summarize, and execute in one flow, you have effectively expanded the trust boundary to match the broadest possible misuse path.
For agentic systems, the risk grows when the model can take delegated actions. Once outputs can trigger workflow steps, data access, or external requests, the issue is no longer just “bad text.” It becomes authorization, least privilege, and blast-radius management. A harmless-seeming prompt can become an operational action if the runtime environment treats it as valid instruction.
That is why the correct control mindset is layered. Secure code still matters, but it must be paired with runtime guardrails, approval points for sensitive actions, and monitoring that can detect unexpected behavior after the model has already accepted a prompt or context object.
Risk and Threat Considerations
LLM risk is driven by trust abuse, not just software defects. Attackers can poison context, smuggle instructions into retrieved content, or induce the model to reveal data and perform actions that were never meant to be available to the requester.
Failure mechanism: The system treats untrusted input or model output as if it were safe enough to influence decisions, tool calls, or data exposure, so the model becomes the path of compromise even when the codebase itself has no exploit.
Impact: The result can be disclosure, unauthorized actions, policy bypass, or business harm without a conventional vulnerability alert, which makes the failure harder to detect through normal application security testing alone.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Agentic AI Top 10 addresses the attack and risk surface, while NIST AI RMF and NIST SP 800-53 Rev 5 set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP Agentic AI Top 10 | ASI03 — Identity & Privilege Abuse | LLM runtime abuse often becomes privilege misuse through delegated actions. |
| ASI02 — Tool Misuse | The core risk is unsafe model-driven tool and action selection. | |
| ASI06 — Memory & Context Poisoning | Prompt injection and manipulated context directly corrupt model decisions. | |
| Recommendation — Constrain agent privileges and require policy checks before tool execution. Restrict tool scope and validate every sensitive tool invocation. Separate trusted instructions from untrusted context and isolate memory writes. | ||
| NIST AI RMF | GenAI Risk Management | GenAI risk here is runtime trust failure, governance, and operational misuse. |
| Recommendation — Establish controls for testing, monitoring, and incident response around model use. | ||
| NIST SP 800-53 Rev 5 | AC-6 — Least Privilege | Model actions should be limited to reduce blast radius from unsafe outputs. |
| AU-6 — Audit Review, Analysis, and Reporting | Runtime misuse needs logs that explain prompts, decisions, and actions. | |
| Recommendation — Limit model and tool permissions to the minimum required for the task. Record model inputs, outputs, and tool calls for review and anomaly detection. | ||
Practitioner Guidance
What to verify: Verify which parts of the LLM workflow are actually trusted. If user text, retrieved documents, or external tool outputs can change system behavior, treat those paths as security-sensitive and review them like an authorization boundary.
What good looks like: The model can summarize and assist, but it cannot silently escalate from interpretation to execution. Sensitive actions should require explicit policy checks, and the system should log enough context to reconstruct why a model produced a risky output or attempted a disallowed action.
Common mistake: Teams often harden the application layer while leaving the prompt, memory, retrieval, and tool chain largely ungoverned. That leaves the most important failure point, the runtime decision layer, only partially controlled.
Practitioner takeaway: If the LLM can influence decisions or actions at runtime, security must be designed around trust boundaries and delegated authority, not just around code quality.
Related resources from NHI Mgmt Group
- Why do deployed LLMs create risk even when the underlying infrastructure is secure?
- Why do APIs create identity risk even when the application code is secure?
- Why do AI coding tools create a security risk even when code looks correct?
- Why do package publishing workflows create supply chain risk even when code reviews exist?