Join our Newsletter — 33% off our NHI Course
Home FAQ Governance, Ownership & Risk How do organisations measure whether phishing training is…
Governance, Ownership & Risk

How do organisations measure whether phishing training is actually changing user behaviour?

← Back to all FAQ
By NHI Mgmt Group Editorial Team Updated September 7, 2026 Domain: Governance, Ownership & Risk

Organisations should track behaviour-based signals, not just completion rates. Useful measures include simulation interactions, reporting activity, follow-up training outcomes, and trend lines in user risk scoring. If those indicators improve over time, the programme is influencing decisions under pressure. If they do not, the training may be compliant on paper but ineffective in practice.

What Organisations Should Measure Instead of Completion Rates

Completion rates show that people opened and finished a module; they do not show whether they changed what they do when a message looks urgent, familiar, or slightly off. Behavioural measurement should focus on whether users pause, verify, report, and avoid unsafe actions under realistic conditions. That is the difference between compliance theatre and a training programme that changes day-to-day decisions. For a control-oriented view of security outcomes, NIST SP 800-53 Rev 5 Security and Privacy Controls is useful because it frames training as part of an observable security control environment, not a one-time administrative event. In practice, many security teams discover the gap only after simulation data and incident reports start telling different stories.

How Behaviour Change Shows Up in Real Programmes

Phishing training changes behaviour when it alters what users do at the moment of decision. The strongest evidence usually comes from repeated, comparable signals over time rather than a single post-training score. That includes how often users click, whether they submit credentials, whether they report the message quickly, and whether they escalate suspicious content through the right channel. The important point is trend direction, not isolated perfection. A programme can still be effective if it reduces risky responses in one population while improving reporting in another.

Good measurement also separates awareness from operational habit. Users may answer quiz questions correctly yet still click a convincing message when they are busy, distracted, or under pressure. That is why organisations should compare training outcomes with simulated phishing results and real-world reporting behaviour. If reporting increases and unsafe interaction decreases, the training is likely shaping decisions. If completion is high but behaviour stays flat, the issue may be message design, audience segmentation, or weak reinforcement rather than user resistance.

  • Track click, submit, and report rates across repeated simulations, not just one campaign.
  • Compare pre-training and post-training behaviour for the same user groups.
  • Look for faster reporting, fewer credential submissions, and lower repeat failure rates.
  • Separate genuine improvement from short-term familiarity with the test format.

The useful question is not whether users can identify phishing in a classroom setting, but whether they behave differently when the message lands in a live inbox. Where the measurement model cannot connect training to observable user action, it stops being a behavioural programme and becomes a content delivery exercise.

When Phishing Metrics Mislead and What That Means for Governance

Tighter measurement often increases administrative overhead, requiring organisations to balance better insight against user friction and programme fatigue. That trade-off matters because some metrics create false confidence. High completion, low click rates on a single template, or a temporary dip after a campaign can all look positive while the underlying risk remains unchanged. The better approach is to treat metrics as evidence of sustained habit change, not proof of immunity.

Edge cases matter. In highly technical or security-aware populations, simulation results can improve quickly and then plateau, so marginal gains may be hard to detect. In high-turnover or contractor-heavy environments, behaviour may vary by cohort more than by training cadence. There is also an industry consensus gap on how much weight to give simulated phishing versus real incident reporting, because each measures a different part of the response chain. Organisations should avoid over-reading one signal in isolation. If people learn the simulation pattern rather than the underlying judgement, the metric improves while the control fails in the real world. That is where the measurement model breaks down.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

MITRE ATT&CK address the attack and risk surface, while CIS Controls v8 and NIST CSF 2.0 set the governance and control requirements practitioners need to meet.

FrameworkControl / ReferenceRelevance
CIS Controls v814 — Security Awareness and Skills TrainingDirectly addresses training effectiveness and user behaviour change.
Recommendation — Measure user behaviour after training and tune campaigns to reduce risky responses.
NIST CSF 2.0PR.AT — Awareness and TrainingMaps training to demonstrable workforce security outcomes, not attendance.
DE.CM — Security Continuous MonitoringSupports ongoing measurement of phishing resilience and response trends.
Recommendation — Track behaviour-based indicators to confirm training changes security decisions. Continuously monitor simulation and reporting trends to validate training impact.
MITRE ATT&CKT1566 — PhishingPhishing simulations and responses align to attacker phishing techniques.
Recommendation — Use phishing technique data to test whether users resist and report suspicious messages.

Practitioner Guidance

What to prioritise: Prioritise measures that show decision quality under pressure, especially repeated simulation outcomes and reporting behaviour. Completion should be treated as a prerequisite, not the outcome.

What to verify: Verify that your metrics distinguish real behavioural change from test familiarity. If users only improve against one lure style or one template, the programme may be teaching pattern recognition rather than safer judgement.

What good looks like: Good programmes show fewer unsafe interactions, more timely reporting, lower repeat failure rates, and improvement that holds across multiple campaign types and user cohorts.

Practitioner takeaway: The most defensible measure is not whether users finished training, but whether they make safer choices consistently when a message tries to pressure them into acting quickly.

Deepen Your Knowledge

Sign up to our weekly newsletter — get 33% off our NHI Foundation Level Course

    NHIMG Editorial Note
    Reviewed and updated by the NHIMG editorial team on September 7, 2026.
    NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org