Look for agents that continue after identifying a suspicious domain, retrieve secrets during email triage, or submit credentials to pages that they later flag as unsafe. Those behaviours show that the governance point is too late in the workflow and that warning logic is not stopping execution.
What warning signs show the control point is too late?
The most useful signal is not a single mistake, but a repeated pattern: the agent recognizes a risky condition after it has already acted on it. If an agent can surface a suspicious domain, a risky prompt, or an unsafe page and still proceed to collect data, the approval or warning layer is not governing execution, it is only annotating it.
That usually means the trust decision is happening after the user action, tool call, or browser interaction that should have been blocked. In practice, the control is weakest when it depends on post hoc review, because the unsafe step has already crossed the trust boundary.
Another sign is inconsistent containment. A healthy agent may warn and stop, or continue only with a clearly bounded exception path. A failing one oscillates between caution and action, which tells you the policy exists but is not being enforced at the moment the agent needs to decide.
What behaviours indicate that an agent is still operating on unsafe trust?
Look for behaviour that shows the agent treats warnings as advisory rather than binding. Examples include retrieving secrets during email triage, following a link it has already classified as suspicious, or continuing a workflow after flagging the destination as unsafe. These are practical signs that the agent can detect risk but not convert that detection into safe refusal.
Another pattern is over-broad trust inheritance. If the agent inherits a human session, a broad token, or a shared workspace permission and then uses that authority across unrelated tasks, the control is failing to separate “can do” from “should do.” That creates a trust gap even when the agent’s own reasoning looks sound.
You should also watch for hidden credential use. If the agent can reach email, browser, ticketing, or SaaS tools and still has access to secrets or session material it does not need for the current task, the problem is usually privilege scope, not model intelligence.
Where should practitioners draw the line between alerting and enforcement?
The line belongs at the point where a single action can create material exposure, not after the fact. If a task can expose credentials, transfer data, or submit a form with account-impacting consequences, the control must decide before execution whether the action is allowed, requires step-up approval, or must be blocked outright.
For agent trust controls, the operational question is whether the guardrail changes the decision or merely records it. Controls that only log suspicious activity are useful for investigation, but they do not prevent misuse. Controls that can narrow scope, interrupt execution, or require fresh authorization are the ones that actually reduce risk.
That is why separation of duties matters even for automated workflows. The component that detects risk should not be the same component that can silently continue the workflow after risk is detected, unless there is a deliberate, reviewable exception path.
Risk and Threat Considerations
When trust controls fail, the risk is usually silent expansion of authority. An agent that can notice danger but still act can become a conduit for secret disclosure, phishing success, or unsafe downstream transactions, especially when it operates at speed across email, browser, and SaaS tools.
Failure mechanism: The agent’s warning logic fires after the decision to act has already been made, or the policy layer lacks the power to stop the action. That leaves a gap between detection and enforcement that attackers, malicious content, or simple workflow drift can exploit.
Impact: The result is not just a missed alert, it is unauthorized execution under trusted credentials, which can lead to credential exposure, data loss, lateral movement, or repeated unsafe actions that look normal in logs.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Agentic AI Top 10, OWASP Non-Human Identity Top 10 and OWASP API Security Top 10 address the attack and risk surface, while NIST Zero Trust (SP 800-207) and NIST CSF 2.0 set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP Agentic AI Top 10 | ASI03 — Identity & Privilege Abuse | Agent trust failures often show up as uncontrolled authority use. |
| ASI02 — Tool Misuse | Unsafe continuation after warnings is a tool-abuse pattern. | |
| ASI09 — Human-Agent Trust Exploitation | Warnings ignored by agents can enable unsafe trust decisions. | |
| Recommendation — Enforce per-action authorization and step-up review when an agent’s authority exceeds the current task. Constrain tool use so flagged actions cannot proceed without fresh approval. Insert confirmation gates where an agent could otherwise exploit implicit trust. | ||
| NIST Zero Trust (SP 800-207) | PR.AA-05 — Least Privilege | Failed trust controls often indicate excess standing authority. |
| PR.PT-01 — Policy Enforcement | The issue is enforcement occurring too late in the workflow. | |
| Recommendation — Remove standing access and scope each agent action to the minimum required privilege. Place policy enforcement before execution so risky actions can be blocked in real time. | ||
| OWASP Non-Human Identity Top 10 | NHI-05 — Overprivileged NHI | Agents that continue after warnings often retain too much privilege. |
| NHI-10 — Human Use of NHI | Unsafe agent behaviour is often caused by humans extending trust too far. | |
| Recommendation — Reduce agent permissions so a warning cannot still lead to high-impact actions. Stop humans from reusing agent access for tasks that need separate authorization. | ||
| NIST CSF 2.0 | PR.AA-01 — Identity Management, Authentication, and Access Control | Agent trust controls depend on access decisions being enforced correctly. |
| DE.CM-01 — The network is monitored to detect potential cybersecurity events | Suspicious agent actions need monitoring signals to reveal failed controls. | |
| Recommendation — Verify that access control decisions are enforced before an agent can act. Monitor agent actions for risky sequences that indicate warnings are being ignored. | ||
| OWASP API Security Top 10 | API2 — Broken Authentication | Credential submission to unsafe pages exposes broken authentication handling. |
| Recommendation — Block flows that let an agent send credentials to untrusted destinations. | ||
Practitioner Guidance
What to verify: Test the workflow end to end and confirm that a risk flag actually blocks the next action, rather than appearing only in audit output. If the agent can still complete the task after a warning, treat that as a control failure.
Decision rule: If the agent can access secrets, send messages, or submit credentials, require an explicit enforcement point before those actions, not a post-action review step. If you cannot prove that enforcement occurs before execution, assume the control is too weak.
What good looks like: Safe systems stop, degrade, or re-request authorization when risk is detected, and they do so consistently across channels. The practitioner takeaway is that agent trust controls succeed only when they bound action in real time; once warning and execution are separated, trust has become a label, not a safeguard.