Use model safety controls for the behaviour the model provider owns, but use agent governance controls for everything that happens after the model is connected to tools and data. The practical boundary is simple: if the risk comes from instructions, tool access, memory, or environment, the deploying organisation owns it and should govern it accordingly.
Where the Boundary Actually Sits
Model safety and agent governance solve different problems, so teams should not treat them as substitutes. Model safety is about the behaviour of the base model as delivered by the provider, while agent governance starts when that model is connected to tools, memory, data, and execution paths that can change real systems. The boundary matters because the risk surface changes once the model can act.
That is why teams should separate the provider-controlled layer from the deployment-controlled layer. If a failure mode depends on tool calls, permissions, context retention, workflow chaining, or environmental access, the question is no longer only about model behaviour. It becomes a question of how the organisation constrains what the agent can do, when it can do it, and how those actions are observed.
For agent-bound systems, the practical unit of control is not the prompt alone, it is the combination of instruction, authorization, and runtime environment. That is also why the agentic AI security guide is the better starting point when the concern is tool misuse, memory poisoning, or runaway automation rather than raw model output quality.
What Changes Once the Model Can Act
Once a model can call tools or read from connected data sources, the most important risks shift from text quality to delegated authority. A safe model can still be deployed unsafely if it is given broad tool access, persistent memory, or an over-permissive environment. In practice, that means the deploying organisation owns the controls that bound the agent’s reach, even when the model itself was supplied by a third party.
This is where identity, authorization, and environment controls become decisive. A team may accept model-provider assurances about refusals, policy tuning, or harmful-content filtering, yet still need to govern the agent’s actual permissions, session scope, and escalation rules. For that reason, the AI Agent Authorisation Guide is directly relevant when the question is who should control tool use, task scope, and approval gates.
Teams should also distinguish static model behaviour from operational drift. If the same agent can behave safely in a sandbox but becomes risky when connected to production data, the boundary has already moved into governance territory. In that case, the right question is not whether the model is “safe enough” in isolation, but whether the surrounding controls prevent harmful actions from reaching systems of record.
How Teams Should Make the Decision in Practice
A simple decision rule works well: if the issue is output quality, harmful content, or model-side refusal behaviour, start with model safety controls. If the issue is what the system can reach, modify, remember, or execute, start with agent governance controls. The latter is usually the more operationally important layer once tools and data are involved, because it determines blast radius even when the model is imperfect.
Teams should map control ownership to the failure point they can actually influence. Model providers can improve baseline behaviour, but they cannot fully govern your tool permissions, connector policies, memory retention, or approval workflow. That is why the AI Agent Observability, Audit and Incident Response Guide matters when you need attribution, logging, and a tested response path for agent actions.
At higher autonomy, the right control set becomes layered: constrain what the model may request, constrain what the agent may do, and verify what it actually did. Teams that only tune prompts often miss the larger risk, which is that tool access and persistence can turn a minor model error into a real-world change. Good governance is therefore not a second best option, it is the control plane that makes deployment safe.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Agentic AI Top 10 addresses the attack surface, NIST AI RMF and NIST SP 800-53 Rev 5 set the technical controls, and ISO/IEC 42001:2023 defines the regulatory obligations.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP Agentic AI Top 10 | ASI03 — Identity & Privilege Abuse | Directly addresses agent authority once tools and permissions are involved. |
| ASI02 — Tool Misuse | Matches the shift from model output to tool-enabled execution risk. | |
| ASI06 — Memory & Context Poisoning | Applies when agent memory or context becomes part of the risk boundary. | |
| Recommendation — Constrain agent authority and require approval gates for sensitive actions. Restrict tool scope and validate each tool call before execution. Isolate memory sources and block sensitive data from persistent context. | ||
| NIST AI RMF | Govern | AI governance is central to deciding which controls belong to deployment oversight. |
| Recommendation — Define accountability for model, agent, and system-level risk ownership. | ||
| ISO/IEC 42001:2023 | AI management system | Fits organisational governance for AI deployment and operating controls. |
| Recommendation — Assign oversight for AI use, monitoring, and change control across deployments. | ||
| NIST SP 800-53 Rev 5 | IA-5 — Authenticator Management | Credential and token lifecycle matter when agents access tools and data. |
| AC-6 — Least Privilege | Least privilege is the core control for limiting agent blast radius. | |
| Recommendation — Manage and rotate the credentials that enable agent access to systems. Limit each agent to the minimum permissions needed for the task. | ||
Practitioner Guidance
What to verify: Before trusting a system, verify where the model ends and the agent begins. If the deployed system can invoke tools, write memory, or act on data, treat those as governance controls under your ownership, not as provider assurances.
Decision rule: If the control would still matter even when the model is replaced by a different vendor, it belongs in agent governance. If the control only changes the model’s generative behaviour, keep it in the model-safety layer.
What good looks like: The system has narrowly scoped permissions, explicit approval points for sensitive actions, and logs that let you reconstruct who or what caused each material action.
Practitioner takeaway: Do not ask whether the model is safe in the abstract, ask whether the deployed agent can do anything harmful with the access it has been given.
Related resources from NHI Mgmt Group
- How should enterprise teams decide between agent-to-tool governance and application-to-model routing in AI platforms?
- What is the difference between human identity governance and AI agent governance?
- How should security teams decide between native ERP controls and a separate governance platform?
- How do identity teams decide whether an AI agent needs a separate governance model?
Deepen Your Knowledge
Free weekly newsletter
Subscribe to the NHI & AI Identity Journal
The latest on NHI and Agentic AI security – articles, research, breaches, news and events every week.
Bonus 33% off our NHI Course when you subscribe.
Reviewed and updated by the NHIMG editorial team on October 8, 2026.
NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org