Look for sessions where private data, untrusted input, and an external communication path all appear in the same task. That combination signals that the agent can be steered from routine work into leakage or misuse. The useful indicator is not model quality, but whether the workflow can support risky action convergence.
What Makes an Agent Workflow Unsafe in Practice
An agent workflow becomes unsafe when the task design lets sensitive context, outside input, and outbound communication meet without enough boundary control. That is when a routine assistant can be redirected into disclosure, misuse, or unintended side effects. Security teams should treat the workflow, not the model, as the unit of analysis.
The first question is whether the workflow can carry data across trust boundaries in a way the operator did not intend. If the same task can read private material, process untrusted instructions, and send results outward, the workflow has enough reach to produce harmful outcomes even if each individual step looks legitimate in isolation.
That is why agent safety is often a composition problem. A task may be harmless while read-only, but once it can combine context gathering, interpretation, and action, the agent inherits a much larger operational blast radius. The unsafe condition is not “the model is confused”; it is “the workflow can converge multiple risky capabilities in one execution path.”
Signals That Risky Action Convergence Is Emerging
The clearest warning sign is repeated co-location of three elements: private data, untrusted input, and an external communication path. Private data can be customer records, internal notes, secrets, or operational context. Untrusted input can arrive through prompts, documents, tickets, chats, attachments, or web content. The external path can be email, API calls, file export, message posting, or any other outbound action.
When those three elements appear together, the workflow can be steered from “analyze” into “act.” Untrusted text can shape the agent’s interpretation of private material, and the external path gives that interpretation somewhere to go. This is the practical indicator security teams should watch for, because it shows the task can cross from assistance into leakage or misuse.
A second signal is when the workflow has unclear separation between reading, deciding, and publishing. If the same step that ingests sensitive context can also prepare an outbound response, the control boundary is thin. Good workflows force a visible break between information intake and any action that could expose or mutate that information.
How Security Teams Should Evaluate and Contain the Workflow
A useful review asks whether the workflow still behaves safely if the untrusted input is hostile. If the answer depends on the model “ignoring bad instructions,” the design is already brittle. Stronger designs limit what the agent can see, what it can send, and what it can do at each stage of the task.
For team use, an Agentic AI Security Guide is most helpful where you need to assess the full attack surface around inputs, memory, tools, and orchestration. For workflows where privilege is the concern, the AI Agent Authorisation Guide helps translate unsafe convergence into concrete limits on what the agent may do per action. If your concern is visibility after deployment, AI Agent Observability, Audit and Incident Response Guide helps teams define what must be logged, attributed, and revoked when the workflow drifts.
Practically, security teams should prioritize boundary questions over benchmark questions. A workflow that performs well on accuracy tests can still be unsafe if it can mix sensitive context with outbound actions. The right control objective is to prevent the agent from being able to assemble a harmful end-to-end path without explicit oversight.
Risk and Threat Considerations
Unsafe agent workflows are risky because they create a direct path from trusted data to untrusted output. Once private material and external communication share the same task context, an attacker only needs to influence one input channel to shape what gets disclosed or executed. That turns ordinary workflow flexibility into a disclosure and misuse problem.
Failure mechanism: The workflow allows untrusted instructions to coexist with sensitive context and outbound capability, so the agent can be induced to transform hidden data into an external action or leak.
Impact: The result can be data exposure, unauthorized actions, fraudulent requests, or downstream misuse that is hard to distinguish from normal agent activity.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Agentic AI Top 10 and OWASP Non-Human Identity Top 10 address the attack and risk surface, while NIST AI RMF, NIST SP 800-53 Rev 5 and NIST Zero Trust (SP 800-207) set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP Agentic AI Top 10 | ASI02 — Tool Misuse | The question centers on workflows that can be steered into unsafe actions through tool use. |
| ASI03 — Identity & Privilege Abuse | Unsafe workflows become material when an agent can use excessive authority across task steps. | |
| ASI09 — Human-Agent Trust Exploitation | The question is about workflows being steered from routine work into misuse through trusted paths. | |
| Recommendation — Restrict agent tools to the minimum action set and block untrusted instructions from reaching tool execution. Bind each action to least privilege and require policy checks before privileged agent operations. Insert confirmation gates where user-facing trust could be abused to trigger harmful agent actions. | ||
| NIST AI RMF | GOVERN — Govern | Agent safety depends on governance of roles, accountability, and oversight for risky workflows. |
| MAP — Map | Mapping the workflow’s inputs, outputs, and dependencies is central to spotting risky action convergence. | |
| MEASURE — Measure | The answer depends on detecting when workflows combine sensitive context, untrusted input, and outbound paths. | |
| Recommendation — Define ownership, oversight, and escalation rules for agent workflows that can touch sensitive data. Map data flows, trust boundaries, and external connections before approving agent deployment. Measure workflow exposure by tracking where private data, untrusted input, and external actions intersect. | ||
| OWASP Non-Human Identity Top 10 | NHI-05 — Overprivileged NHI | Agent workflows become unsafe when they can act with more privilege than the task requires. |
| NHI-10 — Human Use of NHI | Unsafe workflows often emerge when humans route sensitive work through agent credentials or paths. | |
| Recommendation — Reduce standing privilege on agent credentials and remove access that is not task-specific. Prevent humans from using agent identities as convenient proxies for sensitive actions. | ||
| NIST SP 800-53 Rev 5 | AC-6 — Least Privilege | Least privilege directly limits the damage when an agent workflow is steered into misuse. |
| Recommendation — Limit each agent workflow to the minimum permissions needed for the task. | ||
| NIST Zero Trust (SP 800-207) | AC-4 — Information Flow Control | The issue is fundamentally about controlling how information can move through a workflow. |
| Recommendation — Enforce information-flow boundaries between intake, reasoning, and outbound actions. | ||
Practitioner Guidance
What to prioritise: Review workflows that can both read sensitive content and send externally first, because they have the highest chance of risky action convergence. Treat broad tool access plus untrusted inputs as a gating issue, not a tuning issue.
What to verify: Confirm that no workflow can move from intake to outbound action without an explicit policy check, and verify that private context is not silently carried into destinations that were not intended to receive it.
Common mistake: Teams often focus on whether the model answers correctly, then miss the more important question of whether the workflow can be steered into exposing something it should never have been able to send.
Practitioner takeaway: Unsafe agent behavior is usually visible in the workflow shape before it is visible in the output, so boundaries around data flow and action rights matter more than model confidence.