By NHI Mgmt Group Editorial TeamBased on Pillar Security: “The New AI Attack Surface: 3 AI Security Predictions for 2026” (December 3, 2025)

TL;DR: AI security is moving toward inference-time exploitation, indirect injection, poisoned MCP tooling, and agent-to-agent propagation, according to Pillar Security. The governing assumption breaks when data becomes executable and agents can chain trusted inputs into privileged actions without runtime validation.


At a glance

What this is: This is an analysis of three 2026 AI attack vectors that move security focus from code defects to runtime behaviour, including indirect injection, MCP tool poisoning, and agent-to-agent propagation.

Why it matters: IAM, IGA, PAM, and NHI teams need to treat AI agents as runtime actors whose decisions, tool use, and trust boundaries must be governed like access paths, not just software artifacts.

By the numbers:

  • 86% of organizations are blind to AI data flows, having no inventory or visibility into where their AI is connected or what data is exposed, according to IBM Data Breach Report 2025 cited by Pillar Security.
  • 13% of organizations reported breaches involving their AI models or applications, with 97% lacking proper AI access controls, according to IBM Data Breach Report 2025 cited by Pillar Security.

Context

AI agent runtime security is the problem space here: when a model, tool, memory store, or connector can influence the next action at execution time, the attack surface is no longer limited to code defects. Pillar Security frames this as inference-time exploitation, where the system is compromised through data rather than traditional software vulnerabilities.

The identity governance issue is that agents now sit in control loops that can read, decide, and act across data sources and tools. That makes runtime trust, tool authorization, and handoff validation central to agentic AI security, not a secondary implementation detail.

The article is focused on production exposure, not lab theory. Its examples show how a compromised document source, poisoned MCP server, or chained agent communication can turn trusted inputs into privileged actions in live environments.


Key questions

Q: What breaks when AI agents treat data sources as instructions?

A: The boundary between input and action breaks. If a retrieval source, memory store, or tool response can change what the agent does next, malicious content can redirect legitimate automation without exploiting a classic software flaw. That is why runtime validation matters more than static code review in agentic environments.

Q: Why does tool poisoning create such a high-risk access problem for AI agents?

A: Tool poisoning is risky because the model can see more metadata than the user, and it is trained to follow instructions in that metadata as if they were legitimate. In practice, a poisoned tool can steer the agent toward sensitive files, credentials, or other tools, while the user only sees a harmless tool name or summary. That asymmetry makes abuse hard to notice.

Q: What are the signs that an AI security model is failing or becoming unreliable?

A: Common warning signs include rising false positives, missed threats, inconsistent outputs, and recommendations that security teams cannot explain or validate. Poor results often point to weak training data, stale models, or poisoned inputs. When analysts spend more time correcting the model than using it, the system is no longer improving security and may be creating operational noise.

Q: How should teams govern AI agent handoffs across workflows?

A: They should define each handoff as a governed trust boundary, not a casual message exchange. That means assigning ownership, limiting inherited permissions, and validating the source of context before one agent can influence another. Without that discipline, a single compromised agent can contaminate the whole chain.


Technical breakdown

Inference-time exploitation: why data now behaves like code

Inference-time exploitation describes attacks that influence an AI system at runtime by changing what it reads, trusts, or executes. Instead of exploiting a software bug, the attacker manipulates prompts, retrieved content, memory, or tool output so the model treats malicious instruction as legitimate context. This matters because LLMs do not just classify data, they use it to decide what happens next. Once an agent can combine retrieved information with tool access, the boundary between content and control becomes porous. The result is a new attack surface where governance must cover inputs, not just application code.

Practical implication: classify data sources feeding AI agents as execution inputs and validate them before they can influence action.

MCP tool poisoning and trusted integration abuse

Model Context Protocol, or MCP, creates a standardized way for agents to reach tools and data sources. That convenience also creates a trust problem when an agent assumes an MCP server is authoritative because it lives inside the approved tool chain. If a malicious or compromised server can shape recommendations, the agent may incorporate hostile guidance into code, workflows, or access decisions. This is not classic API abuse alone. It is trust inheritance through a tool boundary that the agent treats as safe because it is internal, familiar, and machine-readable.

Practical implication: treat MCP servers as untrusted inputs unless their responses are independently validated before use.

Agent-to-agent propagation and toxic combinations

Agent-to-agent communication becomes dangerous when each agent inherits context and authority from the previous one without strong verification. A message that looks harmless in one system can become a privileged instruction in the next, especially when read access and write access are combined across agents. Pillar Security describes these as toxic combinations because individually safe capabilities can create unsafe outcomes when chained. The core issue is not just compromise of one agent, but propagation through a trust graph where context contamination and privilege inheritance allow the attack to spread beyond its original entry point.

Practical implication: validate inter-agent handoffs as security events and limit any chain that combines untrusted read paths with privileged write paths.


Threat narrative

Attacker objective: The attacker wants to turn trusted runtime inputs into privileged AI actions that bypass normal review and spread through connected systems.

  1. Entry occurs when malicious instructions are embedded in seemingly trusted data sources, compromised MCP responses, or chat messages that an agent consumes as context.
  2. Credentialed or privileged behaviour follows when the agent treats that input as authoritative and uses its existing tool access to act on the attacker’s instruction.
  3. Escalation happens when poisoned outputs move through agent-to-agent handoffs or development workflows, extending the malicious instruction into additional systems and users.
  4. Impact is supply chain manipulation, data exfiltration, or unauthorized code and workflow changes executed through the agent’s legitimate runtime permissions.

Read and download The State of NHI & AI Agent Breach Report 2026, covering 200+ breaches impacting Non-Human Identities including AI Agents.


NHI Mgmt Group analysis

Runtime trust, not code quality, is now the primary control boundary for AI agents. Traditional secure development models assume the threat sits in code, dependencies, or static configuration. This article shows the real break point is runtime behaviour, where data, tools, and memory can all become instruction channels. For identity governance, that means the security question is no longer only who wrote the code, but what the agent was allowed to trust and act on at execution time.

Data has become executable in agentic systems, which collapses the old distinction between input and instruction. That assumption was designed for software that consumed data without treating it as policy. It fails when an AI agent can reinterpret a document, tool response, or chat message as action guidance. The implication is that security architecture must stop assuming retrieval and execution are separable phases.

Agent-to-agent trust graphs create a new identity problem that spans NHI, application control, and governance. A single compromised agent can propagate malicious context through delegated workflows, so the issue is no longer isolated access but compounded authority across handoffs. This is where OWASP-AGENTIC, OWASP-NHI, and NIST-CSF intersect most sharply: the control plane must govern who or what can influence the next decision. Practitioners should treat inter-agent propagation as a first-class governance risk.

Ephemeral authority without validation creates identity blast radius in AI workflows. When agents can fetch, decide, and act across multiple tools in one session, least privilege at provisioning time is no longer enough. The relevant unit of control becomes the runtime action path, not the standing entitlement. This is the point where PAM, NHI governance, and agentic AI identity stop being separate programmes and become one delegated trust problem.

Runtime guardrails must be built around trust boundaries, not around model performance metrics. The article’s scenarios show that better reasoning does not prevent malicious instruction uptake if the agent still trusts a poisoned source. Practitioners should reframe assurance around validation points, source provenance, and handoff controls. The field is moving toward runtime governance because static assurance cannot contain dynamic execution.

From our research library:

What this signals

Runtime guardrails will replace static review as the more important control for agentic systems. Security teams that still focus primarily on code review will miss the point where agents actually make decisions. The practical shift is toward validating prompts, retrieval sources, tool responses, and handoffs before they can alter action.

Agentic AI identity now overlaps with NHI governance in ways most programmes have not modelled. Agents are not just models, and they are not just software. They are runtime actors whose access can expand through tools, context, and delegated workflows, which is why identity teams should manage them as governed non-human actors with explicit boundaries.

Instruction-path control is the emerging design pattern for AI governance. When runtime inputs can become executable, the control objective changes from preventing code defects to constraining what can influence decisions. That is the programme shift practitioners should prepare for across discovery, validation, and handoff governance.


For practitioners

  • Map runtime data paths for agents Identify every source that can alter agent behaviour at inference time, including documents, RAG stores, chat channels, API responses, and memory. Mark each as a potential instruction path, not just a content source.
  • Validate MCP tool responses before agent use Require independent validation for tool outputs that can influence code, access, or workflow decisions. Do not allow internal tool placement alone to confer trust on returned guidance.
  • Separate untrusted read paths from privileged write paths Review agent workflows for combinations where one context source can trigger writes, commits, approvals, or notifications in another system. Remove or isolate chains that let hostile input reach a privileged sink.
  • Inventory agent-to-agent handoffs as governance objects Track which agents pass context, which permissions they inherit, and where validation is missing between them. Treat each handoff as a control point with explicit ownership and review.

Key takeaways

  • AI agent risk is moving from software defects to runtime manipulation, where data, tools, and context can steer privileged actions.
  • The article’s scenarios show how a compromised input source, tool chain, or agent handoff can propagate malicious instructions through legitimate workflows.
  • Governance now has to focus on runtime trust boundaries, because static code controls do not stop executable data from shaping agent behaviour.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

OWASP Agentic AI Top 10 and OWASP Non-Human Identity Top 10 address the attack and risk surface, while NIST CSF 2.0 sets the governance and control requirements practitioners need to meet.

FrameworkControl / ReferenceRelevance
OWASP Agentic AI Top 10ASI02 — Tool MisuseThe article centers on malicious tool responses steering agent actions at runtime.
ASI03 — Identity & Privilege AbuseThe article shows agents inheriting and extending authority across workflows.
Recommendation — Apply ASI02 to validate tool outputs before agents can turn them into actions. Use ASI03 to constrain delegated privileges and stop authority from propagating unchecked.
OWASP Non-Human Identity Top 10NHI-04 — Insecure AuthenticationTrusted integrations and handoffs fail when the agent cannot verify source trust at runtime.
NHI-10 — Human Use of NHIAI agents are being operated as delegated non-human actors with human business impact.
Recommendation — Map agent trust boundaries to NHI-04 and verify the authenticity of each runtime input. Treat agent actions as governed NHI activity and require explicit ownership for each workflow.
NIST CSF 2.0PR.AA-05 — Access Permissions, Entitlements and AuthorizationsThe article is fundamentally about controlling what agents are allowed to do at runtime.
Recommendation — Apply PR.AA-05 to restrict agent entitlements to only the tools and actions they truly need.

Key terms

  • Inference-time exploitation: A compromise pattern where an AI system is manipulated while it is reasoning or acting, rather than through a traditional software vulnerability. The attacker targets the runtime decision process by shaping inputs, tools, or context so the model performs unsafe actions on its own.
  • Indirect injection: A malicious instruction hidden inside data the AI system trusts, such as retrieved documents, tool output, or memory. The system later treats that data as guidance and follows the embedded command, which makes the payload dangerous because the harmful step happens after ingestion.
  • Toxic Risk Combinations: Toxic risk combinations are unsafe interactions between datasets, access permissions, and AI workflows that only become problematic when combined. Individually they may appear harmless, but together they can expose sensitive information, enable re-identification, or create unintended inferences that traditional controls may miss.
  • Runtime Guardrail: A control applied while an AI agent is operating, not just during configuration or review. Guardrails can block dangerous tool calls, require approval for sensitive actions, or stop data leakage before it reaches systems or users.

Deepen your knowledge

NHI governance, agentic AI identity, and machine identity lifecycle are core topics in our NHI Foundation Level course, the industry's only accredited NHI security programme. If you are responsible for identity security strategy or NHI governance in your organisation, it is worth exploring.
NHIMG Editorial Note
Published by the NHIMG editorial team on June 11, 2026.
Updated on October 10, 2026.
NHI Mgmt Group, the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org