Join our Newsletter — 33% off our NHI Course
Home FAQ Cyber Security How should teams include AI activity in security…
Cyber Security

How should teams include AI activity in security awareness benchmarks?

← Back to all FAQ
By NHI Mgmt Group Editorial Team Updated August 21, 2026 Domain: Cyber Security

They should treat AI agents and automated workflows as part of the same risk model when those systems access enterprise data or services. That means tracking their actions, the identities they use, and the access they hold. If machine activity is excluded, the benchmark understates real exposure in hybrid environments.

Why This Matters for Security Teams

Security awareness benchmarks are meant to show whether people and systems behave within acceptable risk thresholds. Once AI agents, scripted automations, and other non-human identities can read data, call APIs, or trigger workflows, excluding their activity creates a false sense of maturity. A benchmark that only measures human behavior misses the access paths most likely to scale quickly and fail quietly.

Current guidance from NIST SP 800-53 Rev 5 Security and Privacy Controls supports treating access, monitoring, and accountability as control objectives, not as human-only concerns. That matters because AI activity can look legitimate while still violating policy, especially when an agent is using approved credentials but operating outside an intended purpose. Security teams that benchmark only phishing clicks, policy attestations, or training completion will miss whether machine-driven access is actually constrained, observed, and reviewed.

Practitioners also need to distinguish between awareness as a training metric and awareness as a governance signal. For AI-enabled environments, benchmarks should indicate whether teams can inventory machine actors, assign ownership, and validate that their actions are visible in logs and investigations. In practice, many security teams encounter AI-related exposure only after an automation has already touched sensitive data, rather than through intentional benchmark design.

How It Works in Practice

The most effective approach is to extend existing awareness benchmarks so they measure both human and machine activity against the same risk themes: visibility, least privilege, acceptable use, and escalation discipline. That does not mean scoring AI like a person. It means including AI activity in the evidence set that proves awareness is working across the environment.

A practical benchmark model usually includes:

  • inventory coverage for AI agents, service accounts, and automated workflows
  • ownership assignment for each non-human identity and its business purpose
  • logging coverage for prompts, tool calls, API requests, and privileged actions where appropriate
  • review cadence for anomalous or high-impact machine actions
  • policy alignment for data access, tool use, and escalation boundaries

This is where identity and AI governance intersect. If an agent uses a persistent credential, the benchmark should reflect whether that credential is scoped, rotated, and monitored like any other privileged access path. If the environment uses retrieval-augmented generation, benchmark evidence should show that the model does not access sources beyond its intended boundary and that output handling is reviewed for leakage risk. For AI-specific control mapping, NIST AI Risk Management Framework is useful for translating governance goals into operational checks, while OWASP Top 10 for Large Language Model Applications helps teams think about prompt injection, data leakage, and tool misuse as measurable exposure points.

Benchmarks should also distinguish between baseline activity and exception handling. A healthy score is not just “the agent ran successfully.” It is “the agent ran within policy, generated auditable evidence, and escalated when it crossed a threshold.” That makes the benchmark useful for security awareness, not just automation reporting. These controls tend to break down when AI systems are deployed through shadow workflows or third-party integrations because ownership, logging, and policy enforcement become fragmented.

Common Variations and Edge Cases

Tighter benchmarking often increases monitoring overhead, requiring organisations to balance better visibility against operational noise and workload. That tradeoff is especially real in environments with many low-risk automations, where over-instrumentation can bury analysts in benign events.

There is no universal standard for how much AI activity must be included in awareness benchmarks yet. Some teams track all machine actions that touch sensitive systems, while others only score high-impact workflows or privileged use cases. The right boundary depends on data sensitivity, regulatory exposure, and whether the AI system can initiate actions independently or only recommend them. For autonomous or semi-autonomous agents, the benchmark should be stricter because intent and execution are no longer separable in the same way they are for human users.

Special cases matter. Shared service accounts can make benchmark evidence misleading if several automations collapse into one identity. Development and test environments may tolerate broader activity, but that should not be confused with production awareness maturity. If an organisation uses AI assistants embedded in collaboration tools, the benchmark should include how users are trained to recognize when the assistant is acting on enterprise data and how those actions are reviewed.

For operational teams, the most important rule is simple: if an AI system can access, transform, or exfiltrate valuable data, its activity belongs in the benchmark. That principle aligns with broader control thinking in NIST SP 800-53 Rev 5 Security and Privacy Controls, even though current guidance suggests the exact reporting model will keep evolving as agentic AI becomes more common.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

OWASP Agentic AI Top 10 and MITRE ATLAS address the attack and risk surface, while NIST AI RMF, NIST CSF 2.0 and NIST AI 600-1 set the governance and control requirements practitioners need to meet.

FrameworkControl / ReferenceRelevance
NIST AI RMFAI RMF maps governance and measurable risk treatment for AI activity.
OWASP Agentic AI Top 10Agentic AI risks shape how autonomous actions should be measured.
NIST CSF 2.0PR.AC-4Least privilege and access governance apply to machine identities too.
NIST AI 600-1GenAI profile helps translate model-specific risks into controls.
MITRE ATLASAML.T0020Adversarial AI threats inform which activities need monitoring and review.

Use AI RMF to define accountable, auditable controls for AI use in benchmarks.

NHIMG Editorial Note
Reviewed and updated by the NHIMG editorial team on August 21, 2026.
NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org