Join our Newsletter — 33% off our NHI Course

How do organisations measure whether phishing training is actually changing user behaviour?

Organisations should track behaviour-based signals, not just completion rates. Useful measures include simulation interactions, reporting activity, follow-up training outcomes, and trend lines in user risk scoring. If those indicators improve over time, the programme is influencing decisions under pressure. If they do not, the training may be compliant on paper but ineffective in practice.

Why This Matters for Security Teams

Phishing training is easy to count and hard to prove. Completion rates, policy acknowledgements, and annual refreshers can all look healthy while users still click, submit credentials, or approve malicious prompts under pressure. That is why measurement has to shift from participation metrics to behaviour change. Current guidance from NIST SP 800-53 Rev 5 Security and Privacy Controls emphasises monitoring, assessment, and continuous improvement rather than one-time awareness events.

For practitioners, the real question is whether training changes how people act when a message is urgent, familiar, or operationally plausible. If users report more suspicious emails, hesitate before acting, and recover faster after simulations, the programme is influencing decisions. If the only improvement is that people finish the module, the organisation has measured attendance, not resilience. This distinction matters because phishing is often the first step in credential theft, business email compromise, and downstream access abuse. In practice, many security teams discover the weakness only after a real lure has already been accepted, rather than through intentional measurement of behaviour.

How It Works in Practice

Behaviour-based measurement starts with baseline data, then compares it over time across multiple signals. A useful programme tracks simulated-phish click rates, credential submission rates, reporting rates, time-to-report, repeat offender trends, and follow-up performance after targeted coaching. On their own, none of these numbers tells the full story. Together, they show whether users are becoming faster at recognising, rejecting, and escalating suspicious messages.

Measurement is stronger when it is tied to workflow. Security teams can combine simulation results with help desk tickets, mailbox report-button usage, and post-incident reviews to see whether training changes actual decisions. The control objective is not simply to reduce clicks; it is to increase safe behaviours under realistic conditions. If users report suspicious messages faster after training, that is a stronger indicator than a lower click rate alone, because it shows active detection rather than passive avoidance.

  • Use a stable baseline before changing content, frequency, or simulation difficulty.
  • Segment results by role, geography, and exposure, since high-risk groups often need different reinforcement.
  • Track leading indicators such as reporting speed and repeated susceptibility, not just lagging indicators.
  • Pair simulation data with short, targeted coaching so the measured change is attributable to the intervention.
  • Review trends quarterly, because one-off spikes can reflect campaign novelty rather than durable behaviour change.

Where possible, map outcomes to NIST Security Awareness and Training controls and reinforce the human side with examples from The State of Secrets in AppSec, which shows how often secure practices break down in real workflows. These controls tend to break down in organisations that run generic simulations without role-specific follow-up, because users learn to recognise the exercise instead of the threat.

Common Variations and Edge Cases

Tighter measurement often increases operational overhead, requiring organisations to balance stronger evidence against privacy, fatigue, and programme cost. A high-frequency simulation schedule can improve data quality, but it can also create desensitisation if every message feels like a test. Current guidance suggests using a mix of low-friction awareness, periodic simulations, and targeted remediation rather than constant punishment-style testing.

There is also no universal standard for interpreting scores. A lower click rate may look positive, but it can be misleading if employees stop interacting with suspicious mail because they are uncertain whom to trust. Likewise, a higher report rate is not automatically better if reports are noisy and overwhelm the security team. The better question is whether users are making safer decisions with less delay and less confusion.

Edge cases matter most in highly regulated or distributed environments. Contractors, multilingual workforces, and frontline staff often need different content and different measurement windows. Organisations should also separate training outcomes from technical controls such as mail filtering, because stronger filtering can reduce visible clicks without changing user behaviour. If a programme cannot explain whether improvement came from training, tooling, or both, the results are too blended to guide policy. For example, phishing lessons learned from CoPhish OAuth Token Theft via Copilot Studio show that user judgement and identity controls both shape outcomes, and training metrics should reflect that reality.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

OWASP Non-Human Identity Top 10 and CSA MAESTRO address the attack and risk surface, while NIST CSF 2.0, NIST SP 800-63 and NIST AI RMF set the governance and control requirements practitioners need to meet.

Framework Control / Reference Relevance
NIST CSF 2.0 PR.AT Training effectiveness is measured through awareness and response behaviour.
NIST SP 800-63 Phishing often targets identity proofing and authentication behaviour.
OWASP Non-Human Identity Top 10 Phishing can compromise credentials and tokens that protect non-human identities.
NIST AI RMF GOVERN Behaviour-based measurement supports governance over human and automated decision risks.
CSA MAESTRO Agentic workflows and tool use raise the stakes of phishing and credential misuse.

Track user reporting and simulation outcomes to show whether awareness efforts change daily decisions.