OWASP LLM testing maps most directly to agentic application risks, while AI governance frameworks matter when the model’s behaviour must be auditable and bounded. For practitioners, the practical decision is to align evaluation evidence with the controls that govern runtime behaviour, delegated access, and retrieval trust.
Which frameworks matter when LLMs use tools and retrieved context?
When an LLM can call tools and consult retrieved context, the framework question shifts from “model quality” to runtime authority, retrieval trust, and delegated action. The frameworks that matter most are the ones that govern agentic risk, permission boundaries, context integrity, and the controls around any API keys, service tokens, or other access material the system uses to act.
Why the Framework Choice Changes Once the Model Can Act
A chat-only model can be evaluated mainly for output quality, safety, and data handling. Once the model can use tools, read indexed content, and trigger side effects, the security question expands to what it is allowed to touch, what it can retrieve, and how much trust the operator places in that retrieved material. That is why agentic application guidance becomes more relevant than generic LLM testing alone.
The practical boundary is whether the model’s decisions can affect real systems, not whether it produces text. Tool use introduces authorization, context poisoning, prompt injection, and excessive agency concerns, while retrieval adds exposure from over-sharing, stale indexing, and permission bypass in the knowledge layer.
For that reason, the most useful framework set is usually a combination of OWASP Agentic AI Top 10, NIST AI Risk Management Framework, and, where the retrieval layer must enforce access boundaries, NIST AI 600-1 GenAI Profile.
Which control families map to tools, retrieval, and delegated access?
Tool use is mainly an authorization and runtime-control problem. Retrieval is mainly an information-governance and trust problem. If the model can issue requests on behalf of a user or workflow, the access model has to be explicit enough that the system can distinguish a helpful action from an unsafe one.
That is where security control catalogs become useful. NIST SP 800-53 Rev. 5 helps map the problem to access control, identification and authentication, auditability, and configuration management. NIST Cybersecurity Framework 2.0 is useful when you want a higher-level governance view of the same issue across identify, protect, detect, respond, and recover.
For AI systems specifically, the most useful interpretation is that retrieval and tools are not just features, they are governed execution paths. If the model can reach a database, ticketing system, email connector, or code execution environment, the framework must address what evidence exists that those paths are constrained, logged, and reviewable.
How to decide whether to emphasize AI governance or appsec-style testing
The deciding factor is whether you are testing the model as a language system or governing it as an action-capable system. If the main concern is hallucination, benchmark quality, or response safety, general AI governance and evaluation frameworks can be sufficient. If the model can take action, the question becomes how to bound that action and validate the permissions it uses.
That is why the best answer is often layered. Use OWASP Agentic AI Top 10 to structure runtime threats such as tool misuse, identity and privilege abuse, and memory or context poisoning. Use NIST AI RMF or NIST AI 600-1 when you need governance evidence, documentation, and risk treatment for the broader GenAI programme.
Risk and Threat Considerations
Tool-enabled LLMs expand the attack surface because retrieved context and delegated actions can be manipulated independently. A model may be safe in isolation yet still leak data, execute an unsafe tool call, or amplify a poisoned instruction once retrieval or external authority is introduced.
Failure mechanism: Attacker-controlled content in retrieved sources, connector data, or prompt context can steer the model into using trusted tools with excessive scope, while weak permission boundaries let a single request reach multiple systems.
Impact: The result can be data exposure, unauthorized side effects, lateral movement through connected systems, or loss of confidence that the model’s actions are bounded and attributable.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Agentic AI Top 10 addresses the attack and risk surface, while NIST AI RMF and NIST SP 800-53 Rev 5 set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP Agentic AI Top 10 | ASI03 — Identity & Privilege Abuse | Tool-enabled LLMs create runtime privilege and delegated-access risk. |
| ASI02 — Tool Misuse | The question centers on model use of tools and external actions. | |
| ASI06 — Memory & Context Poisoning | Retrieved context can be manipulated and alter model behavior. | |
| Recommendation — Enforce least privilege and explicit authorization for every agent action. Restrict tool scope and validate each invocation against policy. Isolate retrieval inputs and screen context for poisoning before use. | ||
| NIST AI RMF | GOVERN — Govern | Tooling and retrieval need AI governance, accountability, and oversight. |
| MAP — Map | The system must identify tool, retrieval, and delegated-access risks. | |
| MANAGE — Manage | The question asks which frameworks guide control of runtime behavior. | |
| Recommendation — Assign ownership, approval, and auditability for action-capable AI systems. Map model interactions, dependencies, and authority boundaries before deployment. Treat tool access, retrieval trust, and logging as managed AI risks. | ||
| NIST SP 800-53 Rev 5 | AC-6 — Least Privilege | Tool calls and connectors should use minimal permissions. |
| AU-2 — Event Logging | Auditable runtime behavior is central when models can act through tools. | |
| IA-5 — Authenticator Management | Access material like API keys and tokens often enables model actions. | |
| Recommendation — Limit every tool, connector, and service account to the minimum needed. Log model prompts, retrieval events, and tool invocations for review. Rotate and protect credentials used by tools, gateways, and retrieval systems. | ||
Practitioner Guidance
What to verify: Verify that every tool call is mediated by an explicit authorization decision, not by prompt text alone. If the model can act on retrieved context, confirm that retrieval respects the caller’s permissions and that sensitive sources are filtered before they reach the context window.
Decision rule: If the model can cause an external change, treat it as an action system first and a language model second. That means you should test delegated authority, context isolation, and tool output handling before you trust output quality metrics.
Practitioner takeaway: The framework choice follows the runtime power of the system, if the model can only talk, test the model; if it can reach tools, data, or actions, govern the whole path from retrieval to authority to auditability.
Related resources from NHI Mgmt Group
Deepen Your Knowledge
Free weekly newsletter
Subscribe to the NHI & AI Identity Journal
The latest on NHI and Agentic AI security – articles, research, breaches, news and events every week.
Bonus 33% off our NHI Course when you subscribe.
Reviewed and updated by the NHIMG editorial team on October 11, 2026.
NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org