Join our Newsletter — 33% off our NHI Course

Why do phishing simulations alone give an incomplete view of user susceptibility?

Phishing simulations capture a moment of user behaviour, but a missed or ignored message can reflect workload, inbox noise, or irrelevance rather than security awareness. That makes click rates an imperfect proxy for knowledge. To understand true susceptibility, organisations need evidence of recognition, judgement, and decision-making across realistic scenarios, not just a single interaction outcome.

What phishing simulations actually measure

Phishing simulations are useful, but they mostly measure one observable outcome: whether a user interacted with a message in a controlled test. That is only a slice of susceptibility. A person may ignore a message because it arrived during a busy period, blended into inbox noise, or seemed irrelevant, not because they lack the judgement to recognise a real threat.

That distinction matters because phishing is not just about clicking. Real susceptibility includes whether a user notices warning signs, pauses before acting, checks context, and chooses a safer path when the request is unusual. A single click result cannot show how a person would behave across different lures, channels, or levels of pressure.

This is why simulation scores should be treated as behavioural snapshots rather than a full measure of awareness. Organisations get a more accurate view when they look for patterns across multiple encounters, such as reporting behaviour, hesitation, validation steps, and the ability to spot deception in realistic scenarios.

Why click rates are an incomplete proxy for awareness

Click rates compress different human responses into one metric. A user who opens, reads, and discards a message after a quick judgement may look identical to a user who never saw it clearly. That can produce false confidence if the organisation assumes low click rates mean strong security understanding.

To understand susceptibility, you need evidence of recognition and decision-making. That includes whether users identify suspicious cues, know when to verify a request, and understand what action to take when something feels off. It also includes the ability to resist social pressure, urgency, and authority cues, which are often more important than raw message recognition.

For higher-fidelity measurement, use scenarios that vary in realism and context. Different phishing themes, different delivery channels, and different follow-up tasks reveal whether users are simply avoiding a test pattern or actually applying judgement. A sound programme measures more than one behaviour so it can separate memorised caution from genuine decision quality.

What a better assessment approach looks like

A stronger approach combines simulations with other evidence. Look at whether users report suspicious messages, whether they escalate to the right channel, whether they verify requests through an independent path, and whether they repeat the same mistake across scenarios. Those signals tell you far more about actual susceptibility than one isolated interaction.

The assessment should also reflect realistic operating conditions. Workload, context switching, mobile usage, and inbox volume all change how people respond. If testing ignores those conditions, it may overstate risk in some cases and understate it in others. The goal is not to punish mistakes, but to understand where judgement breaks down and where the process helps or hinders safe behaviour.

Good measurement therefore distinguishes knowledge from performance. A user may know the right answer in a training module and still fail under time pressure. Conversely, a user may miss one simulation but consistently report suspicious activity or verify requests correctly in real work. The more the programme captures those differences, the more useful it becomes for targeting training and process improvements.

Risk and Threat Considerations

Relying on simulation clicks alone can create a misleading security posture. The main risk is not just underestimating user susceptibility, but also misdirecting training and controls if teams treat a single test outcome as proof of weakness or proof of competence.

Failure mechanism: The metric collapses context, attention, and workload into one binary interaction result, so it cannot separate distraction from poor judgement or show whether users can recognise and verify a real social engineering attempt.

Impact: Organisations may train the wrong behaviours, miss users who need targeted support, and overtrust a metric that does not reflect real-world decision quality or reporting discipline.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

NIST CSF 2.0, CIS Controls v8 and NIST SP 800-53 Rev 5 set the technical controls, while ISO/IEC 27001:2022 defines the regulatory obligations.

Framework Control / Reference Relevance
NIST CSF 2.0 GV.OV-01 — Oversight of Risk Management Click-rate-only metrics need governance oversight to avoid false security conclusions.
Recommendation — Review phishing metrics as part of governance and ensure they measure meaningful user-risk signals.
CIS Controls v8 CIS-14 — Security Awareness and Skills Training Phishing susceptibility is directly addressed through awareness training and measurement.
Recommendation — Measure awareness with reporting and verification behaviours, not click rates alone.
NIST SP 800-53 Rev 5 AT-2 — Awareness Training Training must assess whether users can recognise and respond to social engineering.
AU-6 — Audit Record Review, Analysis, and Reporting Behavioural evidence from simulations and reporting improves interpretation of susceptibility.
Recommendation — Train users on recognition, verification, and reporting behaviours under realistic scenarios. Review simulation and reporting data together to identify repeatable weakness patterns.
ISO/IEC 27001:2022 A.6.3 — Information security awareness, education and training Awareness programmes should build observable security behaviour, not just test click outcomes.
Recommendation — Design awareness activities that validate judgement and escalation, not only recognition.

Practitioner Guidance

What to verify: Treat simulation results as one input only. Verify whether the programme also measures reporting, escalation, verification behaviour, and repeated performance across different lure types and channels.

What good looks like: A mature assessment shows whether users notice suspicious cues, pause before acting, and choose the safer path under realistic conditions. It also distinguishes momentary distraction from a repeatable judgement gap.

Common mistake: Do not equate a low click rate with strong awareness. That can hide weak recognition skills, poor verification habits, or users who simply learn the test pattern.

Practitioner takeaway: Use simulations to observe behaviour, not to define susceptibility on their own. The real question is whether users can recognise, question, and safely respond when the message is believable, urgent, and embedded in normal work.