The main mistake is assuming the boundary between data and instructions still exists. In agentic systems, a webpage, email, or chat message can contain hidden directives that the model may follow. Teams also underestimate memory poisoning and time-shifted prompt injection, where malicious content is stored now and executed later. Sanitization, content isolation, and tool approval gates are essential controls.
When content stops being “data” and starts behaving like an instruction
Teams get this wrong by treating browser text, email text, chat text, and page content as passive inputs. In an agentic workflow, that content can become part of the model’s working context, which means hidden directives, social-engineering language, or tool requests can influence actions unless the system separates untrusted content from executable intent. The control problem is not filtering words alone, it is preserving authority boundaries.
For agents that act in a browser or read from shared collaboration channels, the risk is that the model normalises hostile content into its own plan. That is why browser-driven attack paths and session-bound interactions deserve the same scrutiny as any other privileged execution path, especially when the agent can click, post, send, or approve on a user’s behalf.
One practical way to think about this is to treat browser and computer-use agents as identity-bearing execution surfaces, not as harmless readers, because their access to signed-in sessions changes the blast radius of any injected instruction.
Why hidden instructions work even when the page looks harmless
Prompt injection succeeds because the model often processes the whole payload as one context stream. A page can contain invisible text, a message can include manipulative phrasing, and a document can embed instructions that are only meaningful once the agent has read them alongside its task prompt. The failure is architectural: the system assumes the source of text also defines its trust level.
Time-shifted prompt injection is worse because the malicious instruction does not need to fire immediately. Content can be stored, forwarded, summarised, or resurfaced later, then executed when the agent revisits it under a different task. Memory poisoning follows the same pattern, but the persistence layer becomes the attack vehicle, so a tainted note or saved preference can survive long after the original source is gone.
The operational lesson is to isolate untrusted content from decision-making context. If a workflow must read external text, the agent should receive a constrained representation, not unconstrained raw instructions, and any action that changes state should be gated separately from content ingestion.
That is why memory isolation and write controls for AI agents matter so much: once hostile content is stored, the real bug is no longer the page, it is the persistence of the poisoned state.
What safe handling looks like in agent workflows
Safe handling starts by separating read, reason, and act. Read is where the agent ingests content. Reason is where it interprets the task. Act is where it uses tools, sends messages, or changes records. Those phases should not share the same authority, because a page that is safe to read is not automatically safe to obey.
Good controls therefore combine content sanitization, content isolation, and explicit approval gates for sensitive actions. The approval gate is most important when the requested action would expose data, spend money, send external communication, or alter production systems. In those cases, the question is not whether the content appears trustworthy, but whether the agent has been given enough authority to cause harm if the content is malicious.
Teams should also assume that browser content may be competing with system instructions, rather than merely being processed by them. The safer pattern is to classify sources, strip executable directives from untrusted text, and force the agent to request confirmation before crossing a trust boundary.
Per-action authorisation for AI agents is the right model here, because approval should happen at the point of impact, not at the point of ingestion.
Risk and Threat Considerations
When agents treat content as trustworthy input, attackers gain a path to manipulate downstream actions without needing direct code execution. The highest-risk cases are browser sessions, shared inboxes, collaboration tools, and long-lived memory stores, because they let hostile content persist until the agent reaches a favourable moment to act.
Failure mechanism: The agent blends untrusted text into its instruction set, then follows embedded directives, later-replayed content, or poisoned memory as if they were legitimate task requirements.
Impact: The result can be unauthorized tool use, data disclosure, fraudulent messages, destructive actions, or lateral movement through accounts and connected systems.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Agentic AI Top 10 and MITRE ATT&CK address the attack and risk surface, while NIST AI RMF sets the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP Agentic AI Top 10 | ASI06 — Memory & Context Poisoning | Stored malicious content can later steer agent behavior. |
| ASI03 — Identity & Privilege Abuse | Injected content becomes dangerous when the agent can act with excessive authority. | |
| ASI02 — Tool Misuse | Prompt injection aims to make agents misuse tools or external actions. | |
| Recommendation — Isolate memory writes and review persisted context before reuse. Constrain agent privilege and require policy checks before tool use. Gate sensitive tool calls with approval and explicit policy enforcement. | ||
| NIST AI RMF | GOVERN — Govern | Agent workflows need governance for authority, accountability and oversight. |
| MAP — Map | Teams must identify where untrusted content can affect agent decisions. | |
| MEASURE — Measure | Monitoring is needed to detect poisoned context and unsafe agent behavior. | |
| Recommendation — Assign ownership, approval boundaries and escalation paths for agent actions. Map content sources, memory paths and tool touchpoints before deployment. Measure agent actions, deviations and approval-bypass events continuously. | ||
| MITRE ATT&CK | T1056 — Input Capture | Prompt injection abuses input channels that feed trusted execution paths. |
| T1204 — User Execution | Hidden directives rely on a user or agent acting on delivered content. | |
| Recommendation — Hunt for malicious content that shapes interactive agent inputs. Review interaction points where content can trigger unintended actions. | ||
Practitioner Guidance
What to verify: Verify that every action-capable agent has a hard separation between content ingestion and tool invocation. If a message, page, or note can change state without a second policy decision, the design is still unsafe.
Decision rule: If the content can influence a destructive, external, or irreversible action, require human approval or a policy check at the moment of execution, not just at the moment the content is first read.
Common mistake: Teams often harden prompts but leave the browser session, mailbox, memory store, or tool chain fully trusted. That leaves the agent vulnerable even when the prompt template itself looks polished.
Practitioner takeaway: The security boundary is not “this text came from a browser,” it is “this text was allowed to become an instruction with authority.”
Related resources from NHI Mgmt Group
- What do teams get wrong when they test an AI agent only in the browser or playground?
- What do teams get wrong when they rely on Shadow AI discovery instead of agent-level control?
- What do teams get wrong about AI agent skills when they rely on scanners alone?
- What do teams get wrong when they rely on human-in-the-loop controls for AI?