Text-only controls break because they judge what the agent says, not what it does. A coding agent can refuse in chat and still issue a tool call that reads files, changes state, or reaches external systems. The missing control is execution-time authorization, which must decide whether the action is allowed before the runtime carries it out.
Why text-only safety controls fail for coding agents
Text filters can block a dangerous answer in chat, but they do not govern the next action the agent takes. A coding agent can appear compliant in prose and still invoke a tool, edit a repository, run a command, or reach an external service. That gap matters because the security decision has moved from language to execution, where per-action authorization is what actually constrains impact.
Once an agent has tool access, the relevant question is no longer only "what did it say?" but "what is it allowed to do right now?" That shift explains why text-only controls routinely miss file reads, state changes, secret access, and network calls. In practice, the control point must sit at the tool boundary, not just at the prompt or response layer.
Text-only safety also fails to distinguish harmless language from harmful execution. A model may decline a request verbally, then still produce an allowed-looking sequence that reaches the same outcome through a code editor, shell, package manager, or API client. That is why agent security has to account for runtime authority, not just content moderation. NHIMG’s AI Coding Agents Security Guide and Zero Trust for AI Agents both frame this as an execution-control problem, not a wording problem.
What actually breaks in the workflow
The first failure is blast-radius control. If the agent can read files, write files, or call external systems without a fresh decision at the moment of action, then the chat safety layer is only advisory. That is how a safe-sounding interaction can still lead to repository changes, secret exposure, or destructive commands.
The second failure is intent enforcement. Text controls evaluate the request in isolation, but coding agents act across a chain of steps. A benign-seeming intermediate step can be enough to stage an unsafe outcome later, especially when the agent can reuse context, access cached credentials, or chain tools without reapproval.
The third failure is attribution and reviewability. If the system cannot show which action was approved, by whom, and under what policy, then a refusal in chat provides little operational assurance. A useful agent control plane needs observable decisions, scoped permissions, and revocation paths. NHIMG’s AI Agent Observability, Audit and Incident Response Guide is the natural companion here because it treats actions, not messages, as the audit surface.
Execution-time authorization is the missing control
Execution-time authorization means the agent must be checked at the moment it wants to act, against the specific action, target, and context. That may include whether the task is within scope, whether the file or system is sensitive, whether the request crosses an environment boundary, and whether the agent still has valid standing access. The control should be granular enough to distinguish "can propose" from "can execute".
For coding agents, that usually means task-scoped access, least privilege, and explicit approval gates for risky operations. It also means separating textual persuasion from permission, so the runtime can deny a command even when the model can describe it confidently. The stronger designs treat tool calls like privileged operations, not like ordinary text generation. NHIMG’s AI Agent Authorisation Guide and MCP Security Guide both reinforce that the policy decision belongs at the point of tool use, not after the fact.
External guidance points the same way. The OWASP Agentic AI Top 10 explicitly calls out identity and privilege abuse, while NIST AI Risk Management Framework gives a governance basis for managing action-level risk rather than only output quality.
Risk and Threat Considerations
When only text controls exist, an attacker only needs a path from language to execution. Prompt injection, indirect prompt injection, or poisoned repository content can steer the agent into unsafe tool use even when the chat response looks cautious. The practical risk is not just bad advice, but unauthorized reads, writes, deletions, or outbound calls that occur after the language layer has already passed inspection.
Failure mechanism: The system trusts the agent’s text response as the safety gate, while the actual dangerous step happens in a downstream tool call, shell command, or API request that is never separately authorised.
Impact: Sensitive files, secrets, source code, cloud resources, or production data can be exposed or altered even though the conversation appeared to remain within policy.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Agentic AI Top 10 addresses the attack and risk surface, while NIST AI RMF, NIST SP 800-53 Rev 5 and OWASP ASVS set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP Agentic AI Top 10 | ASI03 — Identity & Privilege Abuse | Coding agents fail when chat safety misses tool-level privilege abuse. |
| Recommendation — Enforce per-action authorization and least privilege for every agent tool call. | ||
| NIST AI RMF | Govern | The issue is AI governance of runtime actions, not just outputs. |
| Recommendation — Set governance so agent actions are approved and monitored at execution time. | ||
| NIST SP 800-53 Rev 5 | IA-9 — Identification and Authentication (Service and Application Accounts) | Agent tool use depends on service-level authentication and bounded access. |
| Recommendation — Authenticate agent-to-service actions and limit credentials to the required scope. | ||
| OWASP ASVS | V8 — Authorization | The core failure is missing authorization on actions, not on text. |
| Recommendation — Verify every privileged action is authorized before it executes. | ||
Practitioner Guidance
What to verify: Check whether every tool, file, shell, and network action has its own policy decision, not just a chat-level filter. If the agent can complete a sensitive workflow after a verbal refusal, the control is not effective.
Decision rule: If an action can change state, access credentials, or reach an external system, require execution-time authorization before the runtime carries it out. Treat read-only inspection differently from write, delete, or export actions.
What good looks like: The agent can explain a task in chat, but it still cannot execute any privileged step unless the policy engine approves that exact action, target, and context.
Practitioner takeaway: Text safety is useful for shaping conversation, but it is not a substitute for runtime control, because the real security boundary is the action the agent is about to perform.
Related resources from NHI Mgmt Group
- Why do coding agents need more than text safety controls?
- What breaks when AI coding agents rely on command allowlists for safety?
- What breaks when coding agents can reach tools and MCP servers without consistent governance and audit controls?
- When is it crucial to implement least-privilege access for AI agents?
Deepen Your Knowledge
Free weekly newsletter
Subscribe to the NHI & AI Identity Journal
The latest on NHI and Agentic AI security – articles, research, breaches, news and events every week.
Bonus 33% off our NHI Course when you subscribe.
Reviewed and updated by the NHIMG editorial team on October 11, 2026.
NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org