They often confuse a narrow behavior metric with overall culture. Phishing click rates matter, but they do not show whether employees report incidents quickly, follow secure workflows, or use safe access habits in daily work. A meaningful assessment combines completion, reporting, access review behavior, and threat context so leaders can distinguish compliance from real defensive participation.
Why This Matters for Security Teams
Phishing simulation scores are easy to report, which is exactly why they are often overused. A low click rate can suggest awareness, but it does not prove that employees will report suspicious activity, verify unusual requests, or resist real-world social engineering under pressure. security culture is broader than a training metric. It includes whether people follow secure workflows, escalate concerns early, and understand how their actions affect identity, access, and incident response.
This matters because leaders sometimes treat a passing score as evidence that behaviour risk has been solved, when it may only show that a test was familiar or narrowly optimised. The NIST Cybersecurity Framework 2.0 places governance and protective practices alongside detection and response, which is a better reminder that culture must be measured across multiple control outcomes, not one awareness exercise. In practice, many security teams encounter weak reporting discipline only after a phishing campaign has already become an account takeover or fraud event.
How It Works in Practice
Organisations get better signal when they treat phishing scores as one input in a wider behaviour model. That model should combine simulation results with reporting rates, time-to-report, secure email handling, MFA adoption, privileged access hygiene, and whether employees follow approved escalation paths when something feels wrong. Current guidance suggests pairing behavioural metrics with operational evidence so leaders can see whether people are actually acting defensively, not merely avoiding a test.
A practical programme usually tracks the full journey from exposure to response:
- Who reported the message, and how quickly they did it.
- Whether the report reached the right queue for triage and containment.
- Whether users reused passwords, approved unexpected MFA prompts, or granted unsafe access.
- Whether teams with higher exposure, such as finance or IT support, show different risk patterns.
- Whether follow-up coaching changes later behaviour, rather than just improving the next simulation score.
For a control-oriented view, the CISA phishing guidance is useful because it emphasises reporting, verification, and response, not just user error counts. Teams should also check whether identity controls support the behaviour they want. If employees are trained to spot suspicious login prompts but shared accounts, weak access review discipline, or inconsistent MFA enrolment remain in place, the culture signal will be misleading. The strongest programmes align phishing outcomes with access governance, incident workflows, and manager accountability. These controls tend to break down when the organisation uses one-off campaigns across very different roles because the metric is too coarse to reflect real operating context.
Common Variations and Edge Cases
Tighter phishing measurement often increases administrative overhead, requiring organisations to balance behavioural insight against the risk of fatigue and gaming. That tradeoff is real: if simulations become predictable or punitive, employees may learn the test rather than improve their judgement. Best practice is evolving toward more contextual measurement, especially where different user groups face different attack surfaces.
There is no universal standard for this yet, but mature programmes adjust for role, privilege, and exposure. For example, executives, finance users, customer support staff, and administrators should not be judged against the same baseline if their email patterns and threat profile differ materially. The same applies to hybrid workforces and high-turnover environments, where a single percentage can hide serious process weakness. Organisations should also be careful not to treat high reporting volume as automatically positive if those reports overwhelm the SOC without clear triage quality.
Identity and access data can sharpen the picture. If users are passing simulations but still approve unexpected requests, ignore session anomalies, or fail to challenge account recovery prompts, the real issue is trust behaviour, not awareness. For that reason, NHI Management Group recommends reading phishing scores alongside access review behaviour, incident reporting quality, and privileged workflow compliance, not as a standalone culture verdict. The OWASP Top 10 for LLM Applications is not a phishing framework, but it reinforces a broader lesson: security measurement only works when it captures how people and systems actually behave under pressure.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
MITRE ATT&CK and OWASP Agentic AI Top 10 address the attack and risk surface, while NIST CSF 2.0, NIST AI RMF and NIST SP 800-63 set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| NIST CSF 2.0 | GV.OC, DE.CM, RS.CO | Culture should be measured through governance, monitoring, and response behaviors. |
| NIST AI RMF | GOVERN | Behaviour metrics need governance to avoid misleading or overfit measurements. |
| NIST SP 800-63 | Phishing resilience depends partly on stronger identity proofing and authentication habits. | |
| MITRE ATT&CK | T1566 | Phishing simulations relate to the same attack pattern used in real intrusions. |
| OWASP Agentic AI Top 10 | Behaviour-only scoring can miss unsafe action paths in human and AI-assisted workflows. |
Reinforce identity and authenticator practices that reduce successful social engineering.
Related resources from NHI Mgmt Group
- What do organisations get wrong when they rely on one-off security testing?
- What do organisations get wrong when they rely on training completion as a security metric?
- What do teams get wrong when they rely on encrypted tunnelling for access security?
- What do organisations get wrong when they treat phishing resistance as a technology project?
Deepen Your Knowledge
Reviewed and updated by the NHIMG editorial team on August 24, 2026.
NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org