Look for command execution that does not match the visible approval, file writes outside the intended workspace, unexpected skill or MCP server installation, and autonomous changes to config or database targets. A second warning sign is when the agent starts resolving names from sources you did not explicitly trust. Those patterns indicate the agent is acting on manipulated context.
What the warning signs actually show
The shift from helpful automation to unsafe execution is usually visible before a full incident. The pattern is not “the model sounds more confident,” but that it begins acting beyond the approval the user can see, the workspace the user intended, or the trust boundary the user set. That is the point where the agent stops being a drafting aid and starts behaving like an execution path.
For ai coding agent, the most important signal is mismatch. If the agent can run commands, create files, install tools, or touch infrastructure without a corresponding human decision, the control plane has drifted out of alignment with the user’s intent. In practice, that often shows up as autonomy plus hidden side effects, especially when the agent starts using trusted developer context as if it were an open-ended permission grant.
Unsafe execution also tends to expand scope quietly. A request to refactor code should not turn into database writes, config changes, package installation, or name resolution from untrusted sources. The AI Coding Agents Security Guide is useful here because it frames the common failure modes around secrets in context, sandbox boundaries, and supply-chain exposure, which are often the first things to break when an agent crosses the line.
What changes when the agent becomes unsafe
The main change is not just more activity, but less reliable attribution of why that activity happened. A safe assistant can recommend or prepare; an unsafe one can execute, persist, or modify state in ways the user did not explicitly approve. Once command execution, file writes, or target selection become decoupled from visible intent, the agent can mutate real systems with a blast radius that is much larger than the original task.
That is why repository trust, MCP trust, and installation trust matter so much. If an agent is allowed to ingest instructions from a repo, pull tools from a server, or resolve names from a source the operator did not explicitly trust, the context itself can become the attack surface. Amazon Q MCP config vulnerability 2026 is a concrete example of how a malicious repository can turn that trust into unauthorized command execution with developer credentials.
Unsafe execution is also visible in the kind of state the agent chooses to touch. Database targets, config files, secrets stores, cloud resources, and CI settings are all high-risk because they convert a coding assistant into a production control path. When those actions appear without a clear approval step, the issue is no longer productivity. It is uncontrolled authority.
Signals that deserve immediate escalation
Look for combinations, not isolated quirks. One unusual command may be a mistake, but repeated command execution outside approval, file writes beyond the intended workspace, or autonomous changes to deploy, auth, or database targets indicate that the agent is following manipulated context rather than explicit instruction. If the agent also starts resolving names, packages, or endpoints from untrusted sources, treat that as a strong sign that the execution path has been redirected.
That same pattern appears in several public attack paths. In Gemini CLI prompt injection flaw 2025, hidden instructions inside a poisoned repository could cause silent command execution and secret exfiltration. In Sentry MCP Agentjacking 2026, a fabricated error path was enough to steer an agent into attacker-chosen code with developer credentials. Those are the same warning signs practitioners should watch for in everyday use, even when no breach has been confirmed.
Risk and Threat Considerations
Once an AI coding agent can execute beyond visible approval, the risk shifts from assistance failure to privilege abuse. The most dangerous condition is not that the model makes a bad suggestion, but that manipulated context causes it to carry out destructive or credentialed actions with a legitimacy the operator never intended.
Failure mechanism: Adversaries or poisoned content steer the agent through trusted channels, such as repository instructions, tool output, or package resolution, so it performs commands, installs tools, or modifies targets outside the approved scope.
Impact: The result can be data loss, credential exposure, unauthorized infrastructure change, supply-chain compromise, or persistent misuse of developer trust that is hard to spot after the fact.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Agentic AI Top 10, MITRE ATT&CK and CSA MAESTRO address the attack and risk surface, while NIST SP 800-53 Rev 5 and NIST Zero Trust (SP 800-207) set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP Agentic AI Top 10 | ASI03 — Identity & Privilege Abuse | Unsafe agent execution centers on unauthorized authority and hidden action scope. |
| ASI02 — Tool Misuse | The warning signs include commands, installs, and tool use beyond intended task scope. | |
| ASI10 — Rogue Agents | An agent crossing into autonomous unsafe execution matches rogue behavior. | |
| Recommendation — Enforce per-action approval and least privilege for any agent step that changes state. Restrict tool invocation to approved actions and block unexpected tool installation. Detect and isolate agents that act outside declared intent or policy. | ||
| MITRE ATT&CK | T1059 — Command and Scripting Interpreter | Unexpected command execution is a core abuse path in unsafe agent behavior. |
| Recommendation — Monitor and constrain shell, script, and interpreter use by agent processes. | ||
| NIST SP 800-53 Rev 5 | AC-6 — Least Privilege | The issue is excessive or unbounded authority to act beyond approval. |
| IA-5 — Authenticator Management | Unsafe execution often abuses tokens, keys, or sessions that the agent can reach. | |
| CM-7 — Least Functionality | Unexpected skill, package, or server installation reflects unnecessary capability expansion. | |
| Recommendation — Limit agent privileges to the minimum required for the assigned task. Rotate, scope, and protect credentials the agent can access or use. Remove unneeded tools, packages, and services from agent execution paths. | ||
| NIST Zero Trust (SP 800-207) | 3.3 — Policy Engine, Policy Administrator, and Policy Enforcement Point | Unsafe execution is a policy enforcement failure for each action the agent takes. |
| Recommendation — Evaluate and enforce policy for every high-impact agent action. | ||
| CSA MAESTRO | GRA — Govern, Risk, and Audit | Agent autonomy needs governance and auditability when actions cross approval boundaries. |
| Recommendation — Define governance, risk, and audit controls for agent execution paths. | ||
Practitioner Guidance
What to verify: The key test is whether every command, file write, and target change can be tied back to a visible user decision. If the agent can act on its own judgment in a way that changes system state, treat that as an authorization problem, not just an AI quality issue.
Decision rule: If the action would be unsafe for a junior operator with broad shell access, it is unsafe for the agent unless the environment tightly constrains scope, destination, and approval. A coding agent should be able to prepare work, but not silently expand its own authority.
What practitioners underestimate: The most serious signal is often not a dramatic destructive action, but a quiet trust shift, such as resolving dependencies, names, or commands from a source that was never approved. That is usually where manipulated context turns into execution.
Practitioner takeaway: Treat unsafe execution as a boundary failure: the agent is no longer just producing content, it is choosing and carrying out actions, so the control objective becomes visible approval, bounded scope, and trustworthy input sources.
Related resources from NHI Mgmt Group
- What are the signs that an agent is crossing from analysis into unsafe execution?
- What is the difference between human identity governance and AI agent governance?
- When does AI agent access create more risk than it reduces?
- What is the difference between governing human access and governing AI agent access?