By NHI Mgmt Group Editorial TeamBased on Zenity: “0Click Attacks: When TTPs Resurface Across Platforms” (September 29, 2025)

TL;DR: AI agent attack techniques are resurfacing across platforms, with Zenity describing how prompt injection, retrieval poisoning, trusted-domain abuse, and image-rendering tricks reappear in Bing, Microsoft 365 Copilot, Salesforce Einstein, and Agentforce. The pattern shows that agent risks are structural, not one-off vulnerabilities, and demand a dedicated security layer rather than vendor-only fixes.


At a glance

What this is: Zenity argues that the same AI agent attack patterns keep reappearing across platforms, showing that prompt injection and related retrieval abuses are an ongoing design problem.

Why it matters: IAM, PAM, and AI governance teams need to treat agent prompt handling, retrieval boundaries, and external content rendering as security controls, not just product features.


Context

Prompt injection in AI agents is the practice of hiding instructions in content that the agent later reads and acts on. In this article, Zenity shows that the problem is not confined to one product or one vendor, because the same technique keeps reappearing wherever agents mix untrusted input with retrieval and external resource rendering.

That matters for identity governance because agents are acting through delegated access, trusted domains, and retrieval contexts that can be manipulated after authorization. The security question is no longer whether an agent can answer a prompt, but whether the surrounding control plane can stop untrusted content from steering privileged behaviour.


Key questions

Q: What breaks when prompt injection reaches an AI agent's retrieval corpus?

A: The boundary between data and instructions breaks. If an agent cannot tell whether retrieved content is a record or a directive, attacker-controlled text can steer decisions long after it was ingested. That makes the retrieval layer part of the trust model, not just a search function.

Q: Why do trusted domains create extra risk for AI agent security?

A: Because trust can be inherited by content that was never meant to deserve it. When an agent treats allowlisted domains as safe inputs, attackers can hide payloads in places the workflow already trusts, then use that trust to trigger exfiltration or manipulation without obvious user action.

Q: How can security teams tell whether an agent is exposed to prompt injection?

A: Look for agents that read untrusted sources, preserve retrieved context across sessions, and can render or call out to external resources automatically. Those are the combinations that make hidden instructions actionable. Exposure rises when the agent can both ingest mixed-trust content and take side effects from it.

Q: How do organisations separate AI governance from AI security testing?

A: AI governance defines what should be allowed, while AI security testing verifies whether the deployed system actually stays within those boundaries. Governance without runtime validation is only policy on paper, especially once agents can retrieve data, call tools, and trigger workflows on their own.


Technical breakdown

How prompt injection survives retrieval and context reuse

Prompt injection works when malicious instructions are embedded in content that an agent later retrieves as if it were trusted context. Retrieval-augmented generation makes this especially dangerous because the model is not only reading user input, it is also consuming prior records, emails, or CRM entries that can carry attacker-controlled instructions. If the agent does not distinguish between data and directives, the retrieved material can override the intended task boundary. That creates a persistent control problem, not a one-time software defect, because the poisoned content remains available until it is detected or removed.

Practical implication: treat retrieved content as untrusted input and separate instruction-bearing context from ordinary records.

Why trusted domains and resource rendering become an exfiltration path

Several of the techniques described by Zenity rely on using a trusted domain or rendering mechanism to make exfiltration look legitimate. If an agent is allowed to embed data in markdown image URLs or similar references, the act of rendering can trigger an outbound request without the user noticing. That is not merely content abuse, it is a boundary failure between agent reasoning and network egress. The attacker does not need to break authentication if they can get the agent to send the data for them through a permitted external fetch.

Practical implication: constrain outbound rendering and network fetch behaviour for agents that process untrusted content.

Why the same agent TTPs keep resurfacing across vendors

The article’s deeper point is that these attacks recur because the underlying primitives are common: untrusted input, retrieval systems, and external resource access. Vendors may patch a specific path, but as long as the platform lets an agent combine those primitives at runtime, a new payload can often recreate the same effect with slight variation. That is why this problem maps to agent security architecture rather than a single vulnerability class. The control gap sits in how agents are authorised, observed, and bounded while they interpret mixed-trust content.

Practical implication: assess agent controls at the architecture layer, not only at the individual vulnerability level.


Threat narrative

Attacker objective: The objective is to make the AI agent leak sensitive data or perform unauthorized actions while appearing to operate normally.

  1. Entry occurs when an attacker plants hidden instructions through a web form or crafted content that later enters an agent's retrieval corpus.
  2. Credential or data access follows when the agent retrieves the poisoned content and treats it as trusted context during a user query.
  3. Escalation happens when the agent uses trusted-domain rendering or markdown image requests to transmit sensitive data outward without obvious user action.
  4. Impact is the exfiltration or manipulation of CRM data, user data, or other sensitive records through the agent's own workflow.

Read and download The State of NHI & AI Agent Breach Report 2026, covering 200+ breaches impacting Non-Human Identities including AI Agents.


NHI Mgmt Group analysis

Prompt injection is now a control-plane problem, not a prompt-quality problem. The article shows that the same abuse pattern can move from one platform to another with only small changes in payload and delivery. That means the relevant failure is not simply bad model output, but the absence of a governance boundary between trusted instructions and untrusted content. Practitioners should treat this as an agent security architecture issue, not a user-training issue.

Trusted-domain abuse is a form of identity and authority laundering. When an agent is allowed to rely on whitelisted domains, it can inherit trust that was never meant for attacker-shaped content. In identity terms, the agent becomes the executor of a request it should never have authenticated in the first place. The implication is that domain trust, content trust, and action trust cannot be the same control.

Retrieval poisoning creates durable identity risk because the malicious instruction outlives the session that introduced it. Once poisoned content enters a CRM or knowledge store, it can influence future agent actions long after the initial injection. That persistence makes the risk structurally closer to standing privilege than to a transient phishing event. Practitioners need to think in terms of polluted context estates, not isolated prompts.

Agentic attack patterns will keep recurring until controls are applied at retrieval, rendering, and action boundaries. Zenity’s examples show that patching one vendor or one product only moves attackers to the next place where agents can combine untrusted input with outbound action. The field needs consistent guardrails for what an agent may read, what it may render, and what it may send. The practical conclusion is that agent governance must become continuous and runtime-aware.

Prompt injection across platforms is becoming the named concept that best captures this class of risk. The useful mental model is not a single vulnerability, but a repeatable technique family that turns normal enterprise content flows into adversarial control paths. That framing helps security teams connect AI agent security, data governance, and egress control in one programme view. Practitioners should measure exposure by where instructions can be stored, retrieved, and acted upon.

What this signals

Prompt injection across platforms is now a repeatable governance pattern. Security teams should stop treating each disclosure as a one-off vendor issue and start mapping where poisoned context can persist across email, CRM, document stores, and agent memory. The security boundary is the retrieval corpus, because that is where attacker intent can survive long enough to be executed later.

Agent security needs controls around what is read, rendered, and sent. If an assistant can ingest untrusted content, render external resources, and emit outbound requests from the same workflow, then the control plane has already been crossed. The operational response is to set distinct trust rules for context, action, and network egress rather than assuming one policy can govern all three.


For practitioners

  • Map agent retrieval boundaries Identify where assistants, copilots, and agentic workflows consume content from email, CRM, files, or tickets and classify those sources as untrusted unless explicitly verified.
  • Restrict rendering and outbound fetches Limit markdown image loading, external content fetching, and other automatic render behaviours for agents that process mixed-trust data.
  • Separate instructions from records Design storage and retrieval layers so that user-generated content cannot silently become executable instructions inside an agent context.
  • Instrument agent egress monitoring Alert on unusual outbound requests, domain changes, and data-bearing URL construction generated by agent actions.
  • Review trusted-domain allowances Reassess allowlisted domains used by agent workflows and remove assumptions that historic trust still applies to current content.

Key takeaways

  • AI agent prompt injection is no longer confined to one platform, because the same abuse pattern keeps reappearing where retrieval and rendering are combined.
  • The practical risk is not just bad outputs. It is data exfiltration, record manipulation, and trust abuse through the agent's own normal workflows.
  • Security teams need runtime controls that separate untrusted content from executable context and restrict agent egress paths.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

OWASP Agentic AI Top 10, OWASP Non-Human Identity Top 10 and MITRE ATT&CK define the specific risk controls and attack patterns relevant to this term.

FrameworkControl / ReferenceRelevance
OWASP Agentic AI Top 10ASI06 — Memory & Context PoisoningThe article centres on poisoned context being reused by agents across platforms.
ASI02 — Tool MisuseThe attacks turn benign agent capabilities into unintended exfiltration actions.
Recommendation — Map retrieval poisoning to ASI06 and isolate untrusted context from agent memory and retrieval layers. Apply ASI02 to constrain which tools and outbound actions an agent can invoke from mixed-trust content.
OWASP Non-Human Identity Top 10NHI-04 — Insecure AuthenticationThe agent inherits trust from domains and contexts it should not authenticate as safe.
NHI-10 — Human Use of NHIHuman-triggered agent workflows are steered into unsafe actions through prompt manipulation.
Recommendation — Use NHI-04 to review how agents authenticate trust in retrieved content and external resources. Apply NHI-10 controls to prevent human-triggered agent sessions from inheriting unsafe content into execution.
MITRE ATT&CKTA0006;TA0010 — Credential Access; ExfiltrationThe attack chain centres on hidden instruction delivery followed by data exfiltration.
Recommendation — Map prompt-injection detections to TA0006 and TA0010 to catch content-driven data theft paths.

Key terms

  • Prompt Injection (Agentic): An attack where malicious instructions are embedded in content that an AI agent reads, causing the agent to execute unintended actions using its own legitimate credentials. A primary vector for agent goal hijacking and identity abuse.
  • Retrieval Poisoning: Retrieval poisoning is the insertion of malicious or misleading content into a data source that an AI system later retrieves. For agents, it creates delayed execution risk because the payload can sit quietly until a user query causes the model to act on it.
  • Trusted-Domain Abuse: A compromise pattern where malicious content is hosted on a legitimate service or platform domain. The domain appears safe, but the content is attacker-controlled, which breaks the assumption that a trusted site always carries trusted instructions or artefacts.
  • Agent Egress Control: Agent egress control is the governance of where an AI agent can send data, which tools it can call, and what external destinations it may reach. It is essential when agents can transform internal information into outbound requests that may leak sensitive content or trigger unintended actions.

Deepen your knowledge

NHI governance, agentic AI identity, and machine identity lifecycle are core topics in our NHI Foundation Level course, the industry's only accredited NHI security programme. If you are building or maturing an identity security programme, it is worth exploring.
NHIMG Editorial Note
Published by the NHIMG editorial team on June 25, 2026.
Updated on October 11, 2026.
NHI Mgmt Group, the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org