Join our Newsletter — 33% off our NHI Course

What are the signs that agent controls are not working?

The clearest warning signs are unexplained tool calls, unexpected privilege use, unreviewed delegated actions, and memory or context changes that alter later decisions. If an agent can move from one system to another without visible policy checks, governance is happening too late. Practitioners should look for behaviour that outpaces their logging and approval boundaries.

What broken agent controls usually look like

When agent controls are failing, the agent is usually doing things the operator did not expect, did not approve, or cannot easily explain. That can include tool use that bypasses the normal request path, actions taken with broader privileges than the task requires, or state changes that shape later behaviour without a corresponding control event. The key signal is not just activity, but activity that outruns governance.

A useful way to read the warning signs is to separate normal autonomy from uncontrolled autonomy. A well-run agent may act quickly, but its actions should still be attributable, bounded, and reviewable. If the system can reach other tools, datasets, or environments without a visible policy decision, then the control boundary is already too thin to trust.

Signs become more obvious when the agent starts accumulating silent dependencies, such as reused sessions, persistent memory that is not reset, or delegated access that is never revalidated. Those patterns often show up first as small inconsistencies: a tool call that should not have been needed, a permission that appears only after the fact, or a decision that changes because earlier context was altered rather than because the underlying task changed.

Which behavioural patterns should you treat as control failures?

Unexplained tool calls are a classic sign because they indicate the agent is selecting or invoking capabilities outside the expected workflow. That does not always mean compromise, but it does mean the policy model is not constraining action the way the design assumed. The same is true when an agent repeatedly reaches for tools that are unnecessary for the task or uses them in an order that no reviewer can justify.

Unexpected privilege use is another strong indicator. If an agent can read, write, deploy, approve, or transfer data in places its task should not touch, the least-privilege assumption has already failed. That failure may come from overbroad delegation, stale tokens, inherited permissions, or an approval path that checks too late to be meaningful.

Memory or context drift is especially important because it changes future decisions without an obvious external event. If the agent starts behaving differently after an injected instruction, a stale note, or a previous run’s retained state, then the control problem is no longer just authorisation. It has become a trust-boundary problem across prompts, memory, and tool selection. The AI Agent Memory Security Guide is useful here because it focuses on poisoning, isolation, and retention discipline.

Another red flag is when delegated actions are unreviewed or untraceable. If a human is supposed to approve high-impact actions, but approvals are skipped, bundled, or rubber-stamped, then the agent is effectively operating with hidden authority. In practice, that is where teams discover that their policy exists only on paper, not in the enforcement path.

How do you tell whether the problem is observability, policy, or delegation?

Start by asking whether the agent is invisible, overpowered, or over-trusted. If the main issue is that you cannot see what happened, then the logging and attribution model is broken. If you can see the action but it should never have been allowed, then policy enforcement or privilege design is broken. If the action was allowed because a human or upstream system delegated too much authority, then the governance model is broken.

That distinction matters because each failure mode has a different fix. Visibility problems call for better audit trails, correlation, and action attribution. Privilege problems call for task-scoped access, short-lived credentials, and per-action checks. Delegation problems call for tighter approval rules and clearer ownership of who can authorise what on behalf of whom. The AI Agent Observability, Audit and Incident Response Guide is a strong reference when you need to separate logging failures from genuine control failure.

It also helps to compare what the agent did with what your approval boundary can actually prove. If a task can complete without a policy decision being logged, then the control is not preventing unsafe action, only detecting it after the fact. At that point, you are depending on review to catch what enforcement should have blocked.

In more mature environments, control failure often shows up as policy mismatch across systems. One system approves the agent, another grants the token, and a third executes the action, but none of them has a full picture of the chain. That is where multi-hop delegation, stale permissions, and hidden privilege accumulation become hard to spot unless you examine the end-to-end path.

Risk and Threat Considerations

Broken agent controls create a fast path from minor behaviour drift to material compromise. The main risk is that a seemingly small autonomy gap, such as overbroad tool access or weak approval enforcement, can expand into unauthorised data movement, privilege escalation, or cross-system action before anyone notices.

Failure mechanism: The control fails when the agent can act outside its intended policy boundary, either because the policy was not enforced at request time, the privilege was too broad, or retained context altered later decisions without review.

Impact: The result is loss of trust in agent output, delayed containment, and a much wider blast radius if the agent can touch multiple systems, credentials, or business workflows before the failure is detected.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

OWASP Agentic AI Top 10 and OWASP Non-Human Identity Top 10 address the attack and risk surface, while NIST SP 800-53 Rev 5 sets the governance and control requirements practitioners need to meet.

Framework Control / Reference Relevance
OWASP Agentic AI Top 10 ASI03 — Identity & Privilege Abuse Agent privilege misuse is central to these warning signs.
ASI06 — Memory & Context Poisoning Context drift and altered memory are explicit failure signals.
Recommendation — Enforce per-action checks to prevent agents from using excessive privilege. Isolate agent memory and validate retained context before reuse.
OWASP Non-Human Identity Top 10 NHI-05 — Overprivileged NHI Unexpected privilege use and standing access are core control failures.
NHI-10 — Human Use of NHI Unreviewed delegated actions show humans are bypassing control boundaries.
Recommendation — Reduce standing access and scope agent credentials to the task. Keep human approval explicit and prevent agents from inheriting unrestricted human access.
NIST SP 800-53 Rev 5 AU-6 — Audit Record Review, Analysis, and Reporting The question hinges on whether agent actions can be observed and explained.
AC-6 — Least Privilege Unexpected privilege use is a direct least-privilege failure mode.
IA-5 — Authenticator Management Delegated access and reused sessions depend on credential lifecycle control.
Recommendation — Review audit records for unexplained tool use and privilege changes. Limit agent permissions to the minimum required for the task. Rotate and expire agent credentials before they become standing access.

Practitioner Guidance

What to verify: Confirm that every high-impact agent action is attributable to a specific request, policy decision, and privilege path. If you cannot reconstruct that chain from logs, treat the control as unproven even if the action was technically successful.

Decision rule: If an agent can perform a sensitive action without a contemporaneous policy check, reclassify that action as standing privilege and reduce the scope until the control is enforced per request. If approvals happen after execution, they are audit records, not control points.

What practitioners underestimate: The most dangerous failures are often not dramatic breaches but slow normalisation of exceptions. Once teams accept “just this once” delegated action, the agent’s actual operating model diverges from the intended one.

Practitioner takeaway: The question is not whether the agent can do useful work, but whether every useful action still leaves a clear, enforceable, and reviewable control trail.