Subscribe to the Non-Human & AI Identity Journal

Notifications
Clear all

AI attackers and deception traps: what defenders need to change


(@nhi-mgmt-group)
Member Moderator
Joined: 1 year ago
Posts: 13010
Topic starter  

TL;DR: Horizon3.ai tested 21 AI models across 10 providers and 47 human red-teamers in 10,962 attacker decisions, finding AI attackers took deception bait more than twice as often and often recognised traps yet attacked them anyway. That widens the case for treating deception as early-warning instrumentation, not a primary misdirection strategy.

NHIMG editorial — based on content published by Horizons.ai: Hacking the Hackers: Can You Still Deceive an AI Attacker?

Questions worth separating out

Q: How should security teams use deception against agentic AI attacks?

A: Security teams should use deception to reshape what an autonomous system believes is real, valuable, and reachable.

Q: Why do AI attackers complicate traditional honeypot strategies?

A: Traditional honeypots assume an attacker will hesitate, misclassify, or avoid suspicious artefacts.

Q: What breaks when deception is used without identity telemetry?

A: Without identity telemetry, deception can generate noise but not clear security decisions.

Practitioner guidance

  • Instrument honeytokens for identity-linked alerting Tie every decoy interaction to user, service account, token, and session context so the alert becomes actionable instead of merely interesting.
  • Reduce the access value of what deception exposes Limit the scope of secrets, API permissions, and workload credentials so a model that touches a trap cannot continue far beyond the baited surface.
  • Pair decoys with automated containment Trigger session review, token revocation, or workload isolation when canary activity appears, because the study suggests recognition does not guarantee disengagement.

What's in the full report

Horizons.ai's full whitepaper covers the operational detail this post intentionally leaves for the source:

  • The per-artifact comparison of AI and human behaviour across file systems, .htaccess files, HTTP responses, and HTTP requests
  • The decision patterns that explain why advanced models recognised traps yet still continued attack behaviour
  • The whitepaper's own framing of how deception should shift from misdirection toward detection in AI-enabled environments
  • The broader implications for teams testing frontier models and self-hosted AI agents

👉 Read Horizons.ai's whitepaper on AI attackers and cyber deception →

AI attackers and deception traps: what defenders need to change?

Explore further

View Full Forum →  |  NHI Foundation Course →



   
Quote
(@mr-nhi)
Member Moderator
Joined: 3 months ago
Posts: 12594
 

Deception has moved from misdirection to telemetry. The study suggests that AI attackers are more likely than humans to touch planted artefacts, which makes deception valuable as an observation layer rather than a standalone diversion tactic. Once a trap is touched, defenders get a high-signal indicator of suspicious activity, but only if that telemetry is connected to identity, session, and workload context. For practitioners, the lesson is to treat canaries as an alert source inside a broader detection program, not as a substitute for prevention.

A question worth separating out:

Q: How do teams know deception is actually reducing risk?

A: Measure whether trap interactions lead to earlier detection, shorter dwell time, and faster containment. If canaries trigger alerts but the attacker still reaches real assets through standing access, the deception layer is functioning as telemetry while the governance layer is still too weak.

👉 Read our full editorial: AI attackers expose the limits of human-era deception tactics



   
ReplyQuote
Share: