Static testing misses the failure modes that emerge only when prompts, retrieval, and tools interact at runtime. That means prompt injection, prompt leakage, and unsafe tool execution can pass code review and still fail in production. Security teams need session-level evaluation because the real attack surface appears after the model assembles context and begins acting.
What static analysis cannot see in LLM security testing
Static analysis can confirm prompt templates, filters, and code paths, but it cannot observe how an LLM behaves once retrieval results, conversation state, and tool outputs are combined at runtime. The gap matters because many failures are emergent: the model only becomes vulnerable after context assembly, which is why offline review often overstates real protection.
That distinction matters most for systems that look safe in a narrow code review but become unsafe when the model can follow instructions from untrusted content, inherit stale context, or execute downstream actions. Runtime evaluation checks the actual control boundary, not just the source code around it.
Runtime failure modes that pass code review
Prompt injection is the clearest example, because the malicious instruction often lives outside the application code and only becomes visible when retrieved text, chat history, or tool output is interpreted by the model. EchoLeak (Microsoft 365 Copilot) 2025 shows why zero-click prompt injection is a runtime problem: the harmful instruction arrived through content the code itself did not flag.
Unsafe tool execution is the second common blind spot. A static scanner can inspect whether a tool is called, but it cannot tell whether the model will overreach, invoke the wrong action, or chain benign tools into an unsafe outcome. That is why runtime testing must include tool permission boundaries, action approval paths, and the model’s actual behavior when adversarial context tries to steer execution.
Context leakage is the third failure mode. Systems that assemble long prompts, retrieval chunks, memory, and hidden instructions can expose information the code review never treated as user-visible. Runtime tests need to verify what the model can quote, infer, forward, or repeat once it has combined all of those inputs into one session.
Why session-level evaluation changes the security verdict
Security teams need session-level evaluation because the attack surface is not a single prompt or a single API call. It is the full interaction sequence, including retrieval, memory, tool selection, and the order in which the model consumes context. That is where prompt leakage, cross-turn contamination, and tool abuse become visible as system behavior instead of isolated code defects.
For agentic and copilot-style systems, runtime evaluation should cover whether a model can be induced to violate role boundaries, reuse stale context, or carry attacker-controlled instructions across turns. Agentic AI Security Guide is useful here because it frames evaluation around inputs, memory, tools, orchestration, and identity rather than only around model output quality.
That same runtime focus also applies to retrieval-heavy designs. Permission-Aware RAG Guide reinforces the point that access control has to be enforced at retrieval time, since a model can only leak what the system makes available to it in the first place. If the retrieval layer is permissive, static prompt review will not save the session.
Risk and Threat Considerations
Static-only testing creates a false sense of security because the most damaging LLM failures are often composition failures, not syntax failures. Attackers exploit that by hiding instructions in retrieved content, chat history, documents, or tool results, then waiting for the model to assemble the attack path at runtime.
Failure mechanism: The system approves prompt patterns and code paths in isolation, but never exercises the live combination of context, retrieval, memory, and tools where the model actually makes decisions.
Impact: Prompt injection, prompt leakage, and unsafe tool execution can survive review, then surface only in production as data exposure, unauthorized actions, or broken trust boundaries.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Agentic AI Top 10, OWASP API Security Top 10 and MITRE ATT&CK define the specific risk controls and attack patterns relevant to this topic.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP Agentic AI Top 10 | ASI03 — Identity & Privilege Abuse | Runtime LLM failures often involve unauthorized tool use and boundary crossing. |
| ASI02 — Tool Misuse | Unsafe tool execution is a core runtime failure mode beyond static code review. | |
| ASI06 — Memory & Context Poisoning | Prompt leakage and contaminated context emerge only when the session runs. | |
| Recommendation — Test agent sessions for privilege escalation and blocked actions before release. Validate tool invocation paths with adversarial session scenarios. Probe memory and context handling for poisoning and cross-turn contamination. | ||
| OWASP API Security Top 10 | API8 — Security Misconfiguration | Runtime trust boundaries for retrieval and tools fail when access is over-permissive. |
| Recommendation — Harden API and tool exposure so the model cannot reach unsafe resources. | ||
| MITRE ATT&CK | T1204 — User Execution | Prompt injection relies on getting a target to act on malicious instructions. |
| Recommendation — Map prompt-injection test cases to adversary execution paths and validate detections. | ||
Practitioner Guidance
What to verify: Test the full session, not just the prompt file. The minimum useful check is whether the model changes behavior when untrusted retrieval text, prior-turn memory, or tool output is present, because that is where real compromise conditions emerge.
Decision rule: If a control only proves that code is syntactically safe, treat it as incomplete for LLM security. Runtime red teaming, adversarial session tests, and tool-abuse scenarios should decide whether a release is fit for production.
Practitioner takeaway: LLM security testing is only credible when it measures the model’s behavior after context is assembled, because that is when the attack surface becomes real.
Related resources from NHI Mgmt Group
Deepen Your Knowledge
Free weekly newsletter
Subscribe to the NHI & AI Identity Journal
The latest on NHI and Agentic AI security – articles, research, breaches, news and events every week.
Bonus 33% off our NHI Course when you subscribe.
Reviewed and updated by the NHIMG editorial team on October 11, 2026.
NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org