The main warning signs are refusal prompts, guardrail fine-tuning, or safety layers that still leave the model free to call tools directly. If the audit record only shows conversation history and not an explicit allow or deny decision, the control boundary is probably still inside the model rather than around it.
What changes when governance sits inside the model boundary?
When governance is inside the model, the model is doing more than reasoning about the task, it is also trying to enforce the decision. That usually means the same runtime that can generate tool calls is also producing the “should I?” answer. The result is softer control, weaker auditability, and a higher chance that the agent can still act even when its textual output sounds cautious.
A clean external boundary looks different: the policy decision is made outside the model, then enforced before any tool call executes. That distinction matters because the strongest signal is not what the model says, but whether an external control point can independently allow, deny, or constrain the request.
How can you tell from the agent's behaviour?
The most visible signs are refusals, warnings, or safety language that do not actually stop action. If the agent can still call tools directly after a “can’t help with that” style response, governance is happening as narration, not as enforcement. Another clue is inconsistent behaviour across similar prompts, where the model appears to negotiate with itself instead of hitting a stable control gate.
Watch for patterns where the model seems to self-police only at the conversation layer. A model may explain policy, cite guardrails, or rewrite the request in safer terms while still retaining operational capability. That is especially common when the agent has broad tool access and the safety mechanism is only shaping output, not constraining execution.
In practice, the distinction becomes clearer when you compare the model transcript with tool telemetry. If you see a tool call that should have required an explicit approval step, or if the audit trail contains only chat turns instead of a discrete permit or deny event, the control boundary is likely still internal.
What proves the control boundary is outside the model?
External governance creates evidence that survives beyond the model’s wording. A strong design has a separate policy engine, a logged decision point, and a tool gateway that can reject actions even if the model attempts them. That is the difference between “the model promised not to act” and “the platform prevented the action.”
The best operational indicator is a traceable sequence: request, policy evaluation, decision, enforcement, then tool execution only if allowed. If the record shows only user prompt, assistant reply, and tool call, the system may still be depending on model behaviour instead of enforcing authorization at the boundary.
For agent governance, that external decision layer is the practical control. NHIMG’s AI Agent Authorisation Guide focuses on per-action policy decisions, delegated authority, and just-in-time access, while the Zero Trust for AI Agents guide shows how to verify the agent, principal, and request before anything executes.
Why does this matter for security and operations?
Inside-the-model governance creates a weak trust boundary. It can reduce unsafe language, but it does not reliably stop tool use, data access, or side effects. That makes it easier for prompt manipulation, policy bypass, or simple model inconsistency to turn a “soft” refusal into a real operational action.
This is why practitioners should treat self-governance as advisory unless it is backed by an independent enforcement layer. The risk is not only malicious abuse, it is also accidental overreach, because the model may be technically capable of acting even when the intended policy says otherwise. The audit trail then becomes misleading, since it records conversation but not authoritative authorization.
See Agentic AI Security Guide for the broader control model around identity, tools, and orchestration, and AI Agent Observability, Audit and Incident Response Guide for what a useful action trail looks like when you need to prove whether the agent was actually constrained.
Risk and Threat Considerations
Inside-model governance is fragile because the same model that is being asked to follow policy is also the component that can be steered, confused, or overruled by prompt injection, tool misuse, or inconsistent safety behaviour. If the policy is not enforced outside the model, a successful prompt attack can turn a refusal into an executed action.
Failure mechanism: The model generates safety language, but the tool layer still accepts the call because no separate policy engine or approval gate exists to stop it.
Impact: Sensitive actions can be executed with no durable authorization record, which increases the chance of unauthorized access, destructive changes, and undetected policy bypass.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Agentic AI Top 10 addresses the attack and risk surface, while NIST AI RMF, NIST Zero Trust (SP 800-207) and NIST SP 800-53 Rev 5 set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP Agentic AI Top 10 | ASI03 — Identity & Privilege Abuse | Model-level governance fails when agent privilege is not enforced externally. |
| ASI02 — Tool Misuse | Direct tool calls after soft refusals indicate tool-use controls are not externalized. | |
| ASI09 — Human-Agent Trust Exploitation | Conversation-only guardrails can create false confidence in safe agent behaviour. | |
| Recommendation — Enforce external per-action policy checks before any agent privilege is used. Gate every tool invocation through a separate authorization decision. Require an auditable decision layer instead of trusting refusal language. | ||
| NIST AI RMF | GOVERN — Govern | Governance inside the model needs organizational oversight, accountability, and policy boundaries. |
| Recommendation — Define and oversee external approval and accountability for agent actions. | ||
| NIST Zero Trust (SP 800-207) | AC-6 — Least Privilege | Agents should not retain direct tool authority when policy is meant to constrain action. |
| Recommendation — Reduce standing agent privilege and enforce least privilege at the control point. | ||
| NIST SP 800-53 Rev 5 | AU-2 — Event Logging | A conversation-only record is insufficient when authorization decisions matter. |
| Recommendation — Log policy decisions and tool executions as separate auditable events. | ||
Practitioner Guidance
What to verify: Confirm that the agent cannot invoke tools unless an external control has emitted an explicit allow decision. If the only record is the conversation, treat the governance boundary as unproven.
Decision rule: If a refusal can be followed by an actual tool call, governance is still advisory. If the platform can block the call independently of model output, the control boundary is where it should be.
What good looks like: Each sensitive action should have a policy decision, an enforcement point, and a log entry that can be reviewed without reconstructing intent from chat text alone.
Practitioner takeaway: The question is not whether the model can speak safely, it is whether something outside the model can stop unsafe action even when the model does not.
Related resources from NHI Mgmt Group
- What are the signs that an AI coding agent is executing host-side commands outside the model tool log?
- What is the difference between human identity governance and AI agent governance?
- When does AI agent access create more risk than it reduces?
- What is the difference between governing human access and governing AI agent access?
Deepen Your Knowledge
Free weekly newsletter
Subscribe to the NHI & AI Identity Journal
The latest on NHI and Agentic AI security – articles, research, breaches, news and events every week.
Bonus 33% off our NHI Course when you subscribe.
Reviewed and updated by the NHIMG editorial team on October 11, 2026.
NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org