Because free text can carry hidden instructions that steer a model toward the wrong action if it is treated as directive content. Security teams should force low-trust input through sanitisation and summarisation before any privileged agent uses it to make a decision.
Why low-trust inputs become dangerous in agentic workflows
Low-trust inputs are dangerous because an agent often treats incoming text as both data and instruction context. If a user message, document, ticket, or web page can influence planning, tool selection, or approvals, hidden instructions can redirect the workflow without ever looking like a command. The risk is not only bad content, but bad content being granted operational authority.
That matters most when the input sits close to a decision boundary. A summariser, router, or planner may pass forward text that appears harmless, yet the downstream agent uses it to choose a tool, escalate a case, retrieve sensitive data, or approve an action. Once that happens, the original low-trust source has effectively crossed from untrusted content into decision influence.
Agentic systems are especially sensitive because the model’s output can trigger follow-on actions. A prompt fragment, attachment note, or chat message can steer the agent toward a wrong branch, and the branch can have real effects if the agent has credentials, API access, or delegated authority. That is why the trust level of the input must be separated from the trust level of the action it may influence.
How prompt injection and instruction smuggling show up in practice
Low-trust input creates risk through instruction smuggling, where the text embeds directives that compete with the system’s intended task. The agent may comply with the most recent, most salient, or most persuasive instruction even when it originated from an untrusted source. In agentic identity workflows, that can become a control failure: the agent acts as if an external party had partial control over its decision logic.
This is particularly problematic when the workflow includes delegation, approval, or access decisions. For example, a request to classify an onboarding ticket should not be able to alter the agent’s interpretation of who owns an identity, which entitlement to grant, or whether a privileged action is justified. If the input can reshape those decisions, the workflow has blurred content handling with authority handling.
Controls such as input sanitisation, summarisation, quoting, and strict separation of instructions from content reduce this exposure. For agent-facing content, agentic AI security should treat untrusted text as a potential control bypass path, not just a quality issue. The operational goal is to preserve the meaning of the request while stripping any ability for the source text to issue commands.
What safe handling looks like before a privileged agent acts
The safest pattern is to force low-trust inputs through a narrowing step before any privileged reasoning. Summarise the content into neutral facts, remove executable or directive language, and make the agent decide only on the sanitised representation. If the workflow needs rich context, keep the raw source available for review, but do not let it directly drive tool use or policy decisions.
This is where trust boundaries matter. A privileged agent should only consume content that has already been classified, reduced, or mediated by a control that understands the source’s risk level. In multi-step flows, the handoff between ingestion, interpretation, and action should be explicit so that a human or control plane can tell which text influenced the final decision. NHIMG’s AI Agent Authorisation Guide is useful here because the core design principle is task-scoped authority, not open-ended obedience to whatever text arrives.
For identity-heavy workflows, the same discipline applies to delegation and privilege. A low-trust prompt should not be able to expand an agent’s access, change its role, or substitute a new principal. If the agent is allowed to take action on behalf of a user, the request content must not be allowed to rewrite that relationship. The Zero Trust for AI Agents model is a good fit because it insists on verifying the principal, the request, and the policy per action.
Risk and Threat Considerations
Low-trust inputs can convert a benign workflow into an attack path when the agent has standing privilege or broad tool access. The most serious failure mode is not incorrect summarisation, but unintended execution, where the agent follows attacker-shaped instructions and performs actions outside the user’s intent.
Failure mechanism: Untrusted text is interpreted as operational guidance, then carried forward into a planner, authoriser, or tool-using agent that lacks strict content isolation.
Impact: The workflow can produce wrong decisions, expose sensitive data, approve unauthorised actions, or trigger downstream compromise through delegated access.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Agentic AI Top 10 addresses the attack and risk surface, while NIST SP 800-53 Rev 5, NIST Zero Trust (SP 800-207) and OWASP ASVS set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP Agentic AI Top 10 | ASI03 — Identity & Privilege Abuse | Low-trust inputs can steer privileged agent decisions and abuse delegated authority. |
| ASI02 — Tool Misuse | Injected instructions can redirect an agent into unsafe or unintended tool actions. | |
| Recommendation — Isolate untrusted input before agent decisions that can change privilege or authority. Constrain tool invocation so only validated intents can trigger actions. | ||
| NIST SP 800-53 Rev 5 | SI-10 — Information Input Validation | Sanitisation and normalisation are central controls for untrusted input in agent workflows. |
| AC-6 — Least Privilege | Privileged agent actions become riskier when input can influence broad access paths. | |
| Recommendation — Validate and sanitise low-trust inputs before they reach decision logic. Limit each agent to the minimum access needed for the task. | ||
| NIST Zero Trust (SP 800-207) | Zero Trust Architecture | Per-action verification and assumed breach reduce the impact of hostile input in agent flows. |
| Recommendation — Verify every request and decision instead of trusting the input source. | ||
| OWASP ASVS | V15 — Secure Coding and Architecture | Agent workflows need architectural separation between untrusted content and executable instructions. |
| Recommendation — Design explicit trust boundaries between content ingestion and privileged execution. | ||
Practitioner Guidance
What to prioritise: Separate content handling from authority handling. If a workflow accepts external text, treat the sanitisation boundary as a control point, not a formatting step.
What to verify: Confirm that the agent cannot read low-trust instructions as policy input, tool directives, or approval conditions. The test is simple: if the text were malicious, could it still influence a privileged action?
Decision rule: If the input can affect access, privilege, routing, or action selection, summarise or normalise it before the agent sees it in decision form. If you cannot confidently remove directive content, do not let that input reach a privileged action path.
Practitioner takeaway: The core control is not banning text, it is ensuring that untrusted text can inform context without ever inheriting decision authority.
Related resources from NHI Mgmt Group
- Why do AI agents create new risk in non-human identity management?
- Why do paper-based education workflows create identity and trust risk?
- Why do identity and telemetry failures create outsized risk in agentic SOC workflows?
- Why do Model Context Protocol calls and similar agentic workflows create governance risk for identity teams?
Deepen Your Knowledge
Free weekly newsletter
Subscribe to the NHI & AI Identity Journal
The latest on NHI and Agentic AI security – articles, research, breaches, news and events every week.
Bonus 33% off our NHI Course when you subscribe.
Reviewed and updated by the NHIMG editorial team on October 7, 2026.
NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org