TL;DR: Horizon3.ai tested 21 AI models across 10 providers and 47 human red-teamers in 10,962 attacker decisions, finding AI attackers took deception bait more than twice as often and often recognised traps yet attacked them anyway. That widens the case for treating deception as early-warning instrumentation, not a primary misdirection strategy.
At a glance
What this is: This whitepaper compares AI attackers and human red-teamers, showing that autonomous models fall for deception more often while also demonstrating a recognition-action gap that makes trap detection and response more complex.
Why it matters: It matters because deception, canaries, and honeytokens now need to support AI-enabled intrusion detection and escalation control, not just human attacker misdirection, especially where AI agents can interact with real credentials, tools, and exposed services.
By the numbers:
- Horizon3.ai tested 21 AI models across 10 providers, analyzing 10,962 attacker decisions and benchmarking their behavior against 47 human red-teamers.
👉 Read Horizons.ai's whitepaper on AI attackers and cyber deception
Context
Cyber deception was designed around human attacker psychology. It assumes an intruder will notice risk, weigh effort against payoff, and respond to obvious traps in predictable ways. That assumption becomes weaker when the attacker is an AI system making rapid decisions across many probes, especially where identity, access, and tool use intersect with autonomous execution.
For identity and security teams, the key issue is not whether deception still has value, but what role it should play in a control stack built for AI-driven intrusion. Honeytokens and canaries can still matter, but they increasingly function as detection and telemetry rather than as a reliable way to redirect an attacker away from real assets.
This is directly relevant to NHI governance because AI attackers often act through compromised credentials, API access, and automated workflows. When the adversary itself is non-human, the defender has to manage both the exposed asset and the speed at which the attacker can test and ignore deception.
Key questions
Q: How should security teams use deception against agentic AI attacks?
A: Security teams should use deception to reshape what an autonomous system believes is real, valuable, and reachable. That means deploying decoys, misleading metadata, and false access paths in places where agentic reconnaissance is likely to start. Deception works best when it is tied to identity telemetry, so teams can see when an attacker is consuming decoys instead of progressing toward privileged systems.
Q: Why do AI attackers complicate traditional honeypot strategies?
A: Traditional honeypots assume an attacker will hesitate, misclassify, or avoid suspicious artefacts. AI attackers can recognise a trap and still continue because task completion can outweigh caution. That makes deception less reliable as a diversion tactic and more useful as proof that the attack path has already reached a sensitive boundary.
Q: What breaks when deception is used without identity telemetry?
A: Without identity telemetry, deception can generate noise but not clear security decisions. Teams may know a decoy was touched, but not whether the same actor is now escalating, pivoting, or probing privileged access. Identity signals turn deception from a standalone trap into a governance control that helps separate curiosity from compromise.
Q: How do teams know deception is actually reducing risk?
A: Measure whether trap interactions lead to earlier detection, shorter dwell time, and faster containment. If canaries trigger alerts but the attacker still reaches real assets through standing access, the deception layer is functioning as telemetry while the governance layer is still too weak.
Technical breakdown
Why AI attackers take deception bait more often than humans
Deception works when an attacker must interpret context, infer risk, and decide whether the suspicious object is worth continuing to investigate. In this study, AI models often accepted planted artifacts such as decoy files or crafted responses because they optimised for task completion rather than human-style caution. That makes the bait itself more visible to the model, but also makes pursuit more aggressive once the model decides the object is relevant. The result is not that AI attackers are naive in a human sense. It is that they are decision engines with different thresholds for uncertainty, and those thresholds can be exploited through structured deception.
Practical implication: treat honeytokens as detection triggers and correlate them with identity and session telemetry.
The recognition-action gap in autonomous attack chains
A recognition-action gap appears when a model identifies a trap or anomalous condition but continues anyway. That can happen because the model is optimising for objective completion, not self-preservation, and because it lacks the intuitive risk aversion humans often apply when something looks wrong. In attacker terms, that means the system can acknowledge a decoy and still probe it, touch it, or exfiltrate from it if the task framing rewards persistence. For defenders, this matters because a trap being recognised does not guarantee the adversary will disengage. Deception therefore becomes a signal of attacker intent rather than a control that reliably diverts the attack path.
Practical implication: build detection logic around trap interaction, not trap avoidance assumptions.
How deception changes when AI agents can use real credentials and tools
Autonomous attackers and AI agents alike may operate through real toolchains, exposed APIs, and credentialed sessions. In that environment, deception can surface malicious activity earlier, but it does not solve the underlying identity problem. If an AI-driven process already has access to secrets, tokens, or service accounts, a decoy may identify the attack while the attacker still retains enough privilege to continue moving. This is why the control conversation must include NHI governance, secret scoping, and privilege boundaries. Deception is strongest when it is embedded inside a broader control plane that constrains what a non-human actor can do after the trap is hit.
Practical implication: pair deception with least privilege, credential segmentation, and rapid revocation paths.
Threat narrative
Attacker objective: The objective is to complete the task path, find reachable real assets, and continue attack execution even after encountering suspicious or deceptive artefacts.
- Entry occurs when an AI-driven attacker or agent interacts with planted infrastructure, decoy files, or crafted HTTP artefacts during reconnaissance and task execution.
- Escalation follows when the model recognises the trap but continues probing, allowing defenders to observe repeated attempts against the same misleading target.
- Impact comes when deception is used as an early-warning mechanism to expose attacker intent before the model reaches real assets, credentials, or privileged workflows.
NHI Mgmt Group analysis
Deception has moved from misdirection to telemetry. The study suggests that AI attackers are more likely than humans to touch planted artefacts, which makes deception valuable as an observation layer rather than a standalone diversion tactic. Once a trap is touched, defenders get a high-signal indicator of suspicious activity, but only if that telemetry is connected to identity, session, and workload context. For practitioners, the lesson is to treat canaries as an alert source inside a broader detection program, not as a substitute for prevention.
The recognition-action gap is a new defensive concept worth naming: trap awareness without trap avoidance. This matters because a model can identify a suspicious decoy and still continue the attack path if task completion is rewarded. That undermines older assumptions that obvious deception will naturally filter out automated threats. In NHI and agentic AI environments, the same pattern can appear when a tool-using system sees a control boundary but still has enough privilege to cross it. Practitioners should assume notice does not equal disengagement.
Identity control determines whether deception becomes useful or merely informative. If an AI attacker reaches decoys through standing credentials, broad API scopes, or persistent service accounts, the trap only reveals the problem after access is already established. That makes NHI governance central to the value of deception, especially where agents or automation can operate with reusable secrets. The practical conclusion is that deception should sit behind strong entitlement boundaries, not in place of them.
Security teams should rethink active defence for non-human adversaries. The strongest implication of the whitepaper is that old assumptions about attacker hesitation no longer hold consistently. AI attackers can be easier to catch in the act, but they may also be more tolerant of suspicious conditions if the objective remains visible. That shifts the programme question from whether deception works to where it belongs in the control stack. For practitioners, the answer is after least privilege, before escalation, and alongside high-fidelity monitoring.
Adaptive deception will become more valuable as autonomous tooling spreads. As AI agents and frontier models become more capable, defenders will need traps that are dynamic, identity-aware, and tied to response automation. Static honeypots will still have value, but only if they are instrumented to tell you who or what touched them and what access path was used. For teams, the priority is to connect deception design to access governance and incident response maturity.
What this signals
Recognition without revocation will become the central failure mode in AI-enabled deception programs. Teams will increasingly see that an attacker or agent can notice a trap, continue execution, and still retain enough access to cause damage. The programme response is to connect deception telemetry to identity governance, session termination, and workload isolation rather than treating decoys as standalone protection.
Trap interaction should now be treated as an access-control signal. In practice, that means canary hits belong in the same operational stream as privileged session anomalies, credential misuse alerts, and unusual tool invocation patterns. For programmes that already struggle to distinguish human from machine activity, the next step is a control model that ties deception to identity-centric response, not just SOC observability.
The stronger the non-human identity footprint becomes, the more important it is to map deception events back to secrets, service accounts, and agent permissions. That alignment is where Top 10 NHI Issues and the NIST AI Risk Management Framework become operationally relevant: not as abstract guidance, but as a way to decide which access paths must be constrained before traps are ever encountered.
For practitioners
- Instrument honeytokens for identity-linked alerting Tie every decoy interaction to user, service account, token, and session context so the alert becomes actionable instead of merely interesting.
- Reduce the access value of what deception exposes Limit the scope of secrets, API permissions, and workload credentials so a model that touches a trap cannot continue far beyond the baited surface.
- Pair decoys with automated containment Trigger session review, token revocation, or workload isolation when canary activity appears, because the study suggests recognition does not guarantee disengagement.
- Use deception as a test of control boundaries Measure whether an AI attacker can reach decoys only after already passing broad entitlements, then tighten the access path that made the trap reachable.
Key takeaways
- AI attackers change the purpose of deception from misdirection to early warning, because they can notice traps and continue anyway.
- The study’s scale, 21 models and 10,962 attacker decisions, shows this is a control-design problem rather than an edge case.
- Defenders should anchor deception inside identity governance, containment, and revocation workflows so trap hits become actionable.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Agentic AI Top 10 and MITRE ATLAS address the attack and risk surface, while NIST AI RMF, NIST CSF 2.0 and NIST SP 800-53 Rev 5 set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP Agentic AI Top 10 | The article concerns agentic attacker behaviour and trap interaction in autonomous systems. | |
| NIST AI RMF | MANAGE | The whitepaper directly affects how organisations govern and monitor AI-driven attack surfaces. |
| MITRE ATLAS | The study addresses adversarial behaviour by AI systems and trap interaction patterns. | |
| NIST CSF 2.0 | DE.CM-1 | Decoy hits and canary interactions are detection signals that fit continuous monitoring. |
| NIST SP 800-53 Rev 5 | SI-4 | Trap interactions should feed security monitoring and automated response controls. |
Assess deception and tool-use risks against agentic AI patterns before allowing access to real systems.
Key terms
- Cyber Deception: Cyber deception is the use of decoys, honeytokens, cloaked assets, and misleading identity signals to make attacker actions harder to validate. In identity security, it changes the environment an intruder sees so that reconnaissance and credential abuse expose intent earlier and reduce the attacker’s ability to trust what they find.
- Recognition-action gap: The distance between understanding a risk and implementing controls that actually reduce it. In AI agent governance, this gap appears when organisations accept that agents are risky but still lack policy enforcement, audit coverage, and revocation mechanisms at runtime.
- Honeytoken: A honeytoken is a deliberately planted secret or credential designed to be detected when used. It helps security teams spot misuse early, especially in environments where machine identities and automation can move faster than manual investigation or containment.
- Non-Human Identity (NHI): A digital identity assigned to a non-human entity such as a software application, service account, API key, bot, machine, or AI agent that enables it to authenticate and interact with systems without direct human involvement. NHIs now outnumber human identities in most enterprises by 25 to 50 times.
What's in the full report
Horizons.ai's full whitepaper covers the operational detail this post intentionally leaves for the source:
- The per-artifact comparison of AI and human behaviour across file systems, .htaccess files, HTTP responses, and HTTP requests
- The decision patterns that explain why advanced models recognised traps yet still continued attack behaviour
- The whitepaper's own framing of how deception should shift from misdirection toward detection in AI-enabled environments
- The broader implications for teams testing frontier models and self-hosted AI agents
Deepen your knowledge
The NHI Foundation Level course, the industry's only accredited NHI security programme, covers NHI governance, secrets management, and workload identity for practitioners who need to control non-human access paths. It is designed for teams building identity-led security programmes across cloud, automation, and agentic AI.
Published by the NHIMG editorial team on August 1, 2026.
NHI Mgmt Group — the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org