Join our Newsletter — 33% off our NHI Course

What do teams get wrong about measuring phishing awareness?

They often measure completion rates or click rates and assume that means risk has improved. Those metrics do not show whether people changed behaviour in real work scenarios. A better measure is whether risky actions decline over time, whether users report suspicious messages faster, and whether high-risk groups receive targeted interventions that reduce repeat exposure.

Why This Matters for Security Teams

phishing awareness is often treated as a training metric, but it is really a behavioural risk signal. Completion rates can show that a programme was delivered, while click rates can show that a test was attempted, yet neither proves that staff are better at making safe decisions under pressure. For that reason, measurement needs to be tied to the security outcome that matters: fewer risky actions, faster reporting, and better containment when a message slips through.

This is where many programmes drift into compliance theatre. If leadership only sees a percentage, the organisation may overestimate maturity and underinvest in controls that actually reduce exposure, such as reporting workflows, mailbox protections, and targeted coaching for repeat-risk groups. The NIST Cybersecurity Framework 2.0 is useful here because it encourages outcome-based thinking across governance, protection, detection, and response rather than vanity metrics alone.

In practice, many security teams discover that phishing “success” was measured so narrowly that the first real indicator of failure was a business email compromise or a credential reset surge, not the awareness dashboard.

How It Works in Practice

A useful phishing measurement model separates training activity from operational control effectiveness. Training completion can still matter for governance, but it should sit below behaviour-based measures that reflect how people act when an email, chat message, or collaboration invite looks suspicious. Good programmes track whether users report suspicious content, whether they escalate it through the correct channel, and whether repeat exposure declines after intervention.

Practitioners usually get better signal when they combine several indicators:

  • reporting time, not just click rate;
  • repeat exposure for the same user or team;
  • quality of reports, including whether the message was genuinely suspicious;
  • incident correlation, such as credential submission, token theft, or mailbox rule abuse;
  • targeted follow-up for roles that face higher external contact volume.

That approach aligns with CISA phishing guidance, which emphasises reporting and response as much as user education. It also fits broader operational resilience thinking in NIST CSF 2.0, because the real objective is to reduce the likelihood that a lure becomes an incident. The best programmes keep test design realistic, vary scenarios by function, and review trends over time instead of chasing a single score.

Measurement also needs context. A finance team that handles invoices, a recruiter that exchanges files with candidates, and an executive assistant that receives calendar invites do not face the same exposure profile. If the programme ignores role-based risk, it will mistake uniform scoring for fairness when it is really just averaging away the highest-risk behaviour. These controls tend to break down in large, distributed organisations with inconsistent reporting channels because local teams interpret the same suspicious message differently and data quality becomes too uneven for reliable trend analysis.

Common Variations and Edge Cases

Tighter phishing measurement often increases programme overhead, requiring organisations to balance richer behavioural insight against privacy, resourcing, and user fatigue. That tradeoff is real, especially where legal teams, unions, or employee councils want strict limits on monitoring and where repeated simulation campaigns can reduce trust if they feel punitive.

Best practice is evolving on how far to go beyond click rates. Some teams now use risk-weighted scoring, but there is no universal standard for this yet, and it can become misleading if the weighting model is not transparent. A low click rate may still hide a weak reporting culture, while a higher click rate in a well-monitored team may be less concerning if suspicious messages are escalated quickly and contained effectively.

Another edge case is automation. If mail filters, browser isolation, and identity controls are strong, awareness metrics may show improvement even when user behaviour has not changed much. That does not make the programme useless, but it means the measure is capturing layered defence rather than awareness alone. For that reason, teams should separate human behaviour metrics from control efficacy metrics and review both. Where phishing feeds directly into credential theft or account takeover, map the programme to response handling as well as awareness so that lessons from incidents feed back into targeted interventions.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

CISA, MITRE ATT&CK and OWASP Agentic AI Top 10 address the attack and risk surface, while NIST CSF 2.0 and NIST SP 800-63 set the governance and control requirements practitioners need to meet.

Framework Control / Reference Relevance
NIST CSF 2.0 GV.OV-01 Outcome-based measurement supports governance oversight of awareness effectiveness.
CISA CISA phishing guidance stresses reporting and response, not just awareness completion.
MITRE ATT&CK T1566 Phishing is the core attack pattern behind user-targeted credential capture.
NIST SP 800-63 IAL2 Credential compromise from phishing affects identity assurance and account recovery.
OWASP Agentic AI Top 10 Agentic workflows can amplify phishing impact through automated actions and trusted prompts.

Review whether automated agents can be tricked into unsafe actions by convincing messages or prompts.