By NHI Mgmt Group Editorial TeamBased on Arkose Labs: “The Agentic AI Security Category Is Converging on the Wrong Answer” (May 1, 2026)

TL;DR: Agentic AI attackers can learn trust boundaries through autonomous iteration, session-to-session learning, and identity spoofing, while the Arkose Labs 2026 Agentic AI Security Report says 97% of enterprise leaders expect a material incident within 12 months but only 6% of security budgets target it. Identity verification alone is fragile when the attacker’s behaviour evolves faster than review cycles can respond.


At a glance

What this is: This is Arkose Labs' argument that agentic AI security is over-indexing on identity verification while attackers learn trust boundaries through repeated, autonomous interaction.

Why it matters: It matters because IAM and security teams cannot govern agentic traffic safely if they only know who the agent claims to be and never measure what it does at runtime.


Context

Agentic AI security is becoming a governance problem, not just a detection problem. The core issue is that current trust-layer designs assume identity verification is enough to separate legitimate automation from harmful agent behaviour.

That assumption breaks down when an agent can probe, adapt, and iterate through sessions faster than review cycles can react. For identity and access teams, the real question is whether policy is being enforced at the point of behaviour, not just at the point of authentication.


Key questions

Q: What breaks when identity verification and authorization are handled separately for AI agents?

A: Trust becomes uneven. Strong proofing may confirm who or what is entering the environment, but it does not stop an overreaching agent once access is granted. Runtime authorization without reliable proofing also risks enforcing policy against the wrong actor or delegation context.

Q: Why do agentic AI attacks get past trust layers over time?

A: Because the attacker can probe repeatedly, vary timing and credential patterns, and learn the boundary from the responses. The trust model becomes a training surface. Once the decision boundary is legible, the attacker only needs to emulate the expected pattern well enough to pass.

Q: How do organisations know if agentic AI governance is actually working?

A: Look for three signals: access decisions tied to task context, complete audit records linking agents to datasets, and rapid revocation when scope changes. If reviewers still need manual reconstruction after an incident, the programme is not mature. Effective governance produces explainable access, not just allowed or denied results.

Q: When is economic deterrence more effective than trust classification for agents?

A: It becomes more effective when attackers can learn classification rules faster than teams can refresh them. In that case, the objective is not perfect detection but making each probe expensive enough that repeated exploitation becomes uneconomic. The control question changes from 'who is it?' to 'what does it cost to keep trying?'


Technical breakdown

Why agent identity is not the same as agent behaviour

Identity verification answers a narrow question: who or what is interacting with the system. Behavioural security answers a different one: what is the entity trying to do, and how does it adapt when blocked. In agentic AI environments, those two signals diverge quickly because an attacker can present a plausible identity while using autonomous iteration to learn the decision boundary. That makes identity a useful input, but not a complete control model. Security teams that stop at classification are measuring provenance, not intent or execution pattern. The result is a control plane that can authenticate an agent and still fail to recognise harmful runtime behaviour.

Practical implication: treat identity as an input to policy, not as proof of safe behaviour.

How autonomous probing defeats trust layers

A trust layer can only be as strong as the assumptions baked into its boundary. Agentic attackers exploit this by running many legitimate-looking sessions, varying timing, credential patterns and interaction cues until the model's edges become visible. That is not a single bypass event. It is a learning process that turns the control itself into a target. Once the boundary is legible, the attacker no longer needs to break verification. They only need to look enough like an authorised agent to pass. This is why static trust frameworks degrade against machine-speed adaptation: the boundary is discoverable, and discovery becomes exploitation.

Practical implication: assume your classification boundary will be probed until it is learned.

Why the interaction layer is the only durable control point

The interaction layer is where an agent actually creates account events, submits forms, calls APIs and completes transactions. That is where behavioural evidence exists, because it is the point where intent meets action. Network-layer telemetry can show source, timing and transport patterns, but it cannot reliably distinguish helpful automation from fraud or policy abuse once the session is underway. Interaction-layer controls can, because they generate challenge-response signals, friction costs and observable deviations from expected execution patterns. The architectural lesson is simple: if the defence only sees the outer shape of traffic, it will miss the meaning of the interaction.

Practical implication: place enforceable controls at the transaction layer where agent behaviour becomes visible.


Threat narrative

Attacker objective: The attacker wants to normalise harmful agent traffic so it can complete fraud or misuse without being stopped by identity-only controls.

  1. Entry begins with sessions that appear legitimate enough to pass identity checks and start interacting with the platform.
  2. Escalation occurs as the attacker varies timing, credentials and behavioural signals to learn the trust boundary across repeated sessions.
  3. Impact follows when the attacker can reliably imitate authorised behaviour and use that learned boundary to sustain fraud or misuse at machine speed.

Read and download The State of NHI & AI Agent Breach Report 2026, covering 200+ breaches impacting Non-Human Identities including AI Agents.


NHI Mgmt Group analysis

Identity verification is not a control model for agentic AI. It is a provenance check, and provenance does not tell you whether an agent is acting within legitimate behavioural bounds. The industry is at risk of hardening around a control plane that can answer who the agent is while remaining blind to what the agent is doing. Practitioners need to treat identity as a necessary input, not the security outcome.

Agentic AI security fails when it assumes the decision boundary is stable. Autonomous probing turns the trust layer into a training surface, because the attacker can iterate until the model becomes predictable. That means the real governance gap is not missing verification alone. It is the assumption that classification stays useful once an adversary can learn from the interaction itself. The implication is that policy must hold even when the boundary is known.

Economic deterrence is the named concept that better fits this threat than trust layering. The article's core argument is that attacks become less viable when every probe increases cost faster than the attacker can amortise it. That reframes agentic security from authentication-centric control to cost-centric governance. For practitioners, the lesson is that durable defence must change the attacker economics, not just improve the confidence of the trust decision.

Interaction-layer governance is becoming the missing control plane for agentic traffic. Security teams need visibility into which agents are active, which flows they touch and how behaviour changes over time. Without that layer, organisations will continue to treat all automation as equivalent, even when some of it is adversarial. The governance task is to make behaviour measurable where the transaction actually happens.

The category will keep repeating the same mistake if it confuses agent identity with safe delegation. A system can verify that an agent is real and still fail to govern whether that agent should be trusted in a particular flow. That is why access policy for agentic systems must be evaluated at the point of action, not only at onboarding or enrollment. Practitioners should redesign controls around observed behaviour, not claimed identity.

From our research library:

What this signals

Economic deterrence changes the governance target: once agentic traffic can learn a boundary, teams have to make repeated probing expensive rather than assuming a single trust decision will hold. That shifts programme design toward transaction-level friction, not just identity proofing.

Only 13% of organisations feel extremely prepared for the reality of agentic AI despite the majority racing toward autonomous adoption, according to the 2026 Infrastructure Identity Survey. That gap means many teams are still building controls for the wrong failure mode.


For practitioners

  • Separate agent identity from behavioural authorisation Map which decisions are made from identity claims alone and which require runtime behavioural evidence. Use that inventory to identify flows where trust layers are acting as the only gate.
  • Instrument the interaction layer Collect signals from account creation, login, checkout and API workflows so agent traffic is measured where actions occur. Behavioural controls are only useful if the transaction layer emits enforceable evidence.
  • Define policy by agent type and risk Create separate handling for authorised agents, adversarial automation and ambiguous gray-area automation, with different thresholds for allow, monitor, challenge and block decisions.
  • Stress-test trust boundaries with repeated probing Use red-team exercises that vary timing, credential patterns and session behaviour to see how quickly a model boundary becomes legible. The goal is to find where the trust layer can be learned rather than breached.

Key takeaways

  • Identity-first controls are insufficient when agentic systems can learn the policy boundary through repeated interaction and adapt around it.
  • The article's central evidence is a governance mismatch, not a detection gap: trust layers can classify an agent while missing the behaviour that makes it dangerous.
  • Behavioural controls at the interaction layer are the decisive shift, because they make attack cost the variable that security teams can actually influence.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

OWASP Agentic AI Top 10 and OWASP Non-Human Identity Top 10 address the attack and risk surface, while NIST AI RMF and NIST CSF 2.0 set the governance and control requirements practitioners need to meet.

FrameworkControl / ReferenceRelevance
OWASP Agentic AI Top 10ASI03 — Identity & Privilege AbuseThe article centres on agents being trusted by identity while behaving outside authorised bounds.
ASI01 — Agent Goal HijackAutonomous probing is used to learn and steer the system toward unintended outcomes.
Recommendation — Map agent identity trust failures to ASI03 and require runtime behaviour checks before granting action scope. Assess whether agent goals can be redirected through repeated interaction and add guardrails where they can.
OWASP Non-Human Identity Top 10NHI-04 — Insecure AuthenticationThe post argues that authenticating an agent does not guarantee safe behaviour in the session.
NHI-10 — Human Use of NHIThe article shows how real agents can be mistaken for safe automation or trusted identities.
Recommendation — Use NHI-04 to separate successful authentication from trustworthy runtime execution. Apply NHI-10 controls to ensure agent use is explicitly authorised and continuously bounded.
NIST AI RMFGOVERN — AI Governance and AccountabilityThe piece is fundamentally about governance assumptions for AI systems and who owns behavioural controls.
Recommendation — Establish governance that assigns accountability for agent behaviour, not just agent onboarding.
NIST CSF 2.0PR.AA-05 — Access Permissions, Entitlements and AuthorizationsThe article shows why authorisation must be enforced against observed behaviour, not only identity claims.
Recommendation — Review authorisation decisions so they depend on context and runtime signals, not just verified identity.

Key terms

  • Agentic AI Security: Agentic AI security is the discipline of securing autonomous AI systems that can take actions, use tools, and chain decisions without direct human approval at each step. It covers identity and access management for AI agents, prompt injection defence, tool call governance, credential scoping, and runtime monitoring. As agentic systems acquire real-world authority, API access, file writes, workflow triggers, the security model must treat them as non-human identities with explicit lifecycle controls, not trusted processes.
  • Interaction layer: The point where a user, agent, or automation interacts with the business flow, such as login, checkout, account creation, or API use. This layer matters because it exposes behaviour, not just network characteristics, and it is often where agentic misuse becomes visible before deeper compromise occurs.
  • Economic Deterrence: A control strategy that makes abuse too costly to sustain. Rather than relying only on detection and blocking, it increases attacker time, effort, and compute until the expected return from targeting a system becomes unattractive.
  • Behavioural Authorisation: Authorisation that evaluates what an actor is trying to do, not only who or what it is. For autonomous or agentic systems, this adds context such as task scope, timing, data sensitivity and safety conditions to standard identity-based controls.

Deepen your knowledge

NHI governance, agentic AI identity, and machine identity lifecycle are core topics in our NHI Foundation Level course, the industry's only accredited NHI security programme. If you are building or maturing an IAM programme, it is worth exploring.
NHIMG Editorial Note
Published by the NHIMG editorial team on June 9, 2026.
Updated on October 10, 2026.
NHI Mgmt Group, the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org