Join our Newsletter — 33% off our NHI Course

What are the signs that a phishing awareness program is actually improving resilience?

A stronger program should show fewer users falling for simulations, more users reporting suspicious messages, and a better resilience factor over time. The report notes that resilience factor improved from 1.5 in 2021 to 2.0 in 2023, which indicates better reporting relative to failure. Tracking both reporting and failure rates gives a clearer picture than training attendance alone.

How to tell whether the program is changing behaviour, not just completing training

The clearest sign of improvement is a shift in user behaviour. If awareness is working, fewer people should click or submit credentials in simulations, and more people should pause, verify, and report suspicious messages before harm spreads. That behavioural change matters more than attendance or course completion, because resilience depends on response quality at the moment the message arrives.

A useful way to judge the program is to compare failure rate and reporting rate together. A lower failure rate without a higher reporting rate may mean people are simply getting better at spotting simulations, while real messages still go unreported. A stronger program changes both outcomes: fewer false actions and faster escalation of suspicious activity.

The strongest indicator is trend direction over multiple cycles, not a single campaign result. If reporting improves steadily while susceptibility declines, the program is likely building habitual resilience rather than producing a short-lived test effect. That is why trend lines are more useful than one-off pass or fail counts.

What a resilience factor actually tells you

Resilience factor is useful because it shows the relationship between reporting and failure, not just raw click rates. In practical terms, it answers whether the organisation is becoming better at turning risky user encounters into defensive action. A rising ratio means the workforce is more likely to surface suspicious mail than to fall for it, which is the behaviour you want.

Used well, the metric helps avoid a common measurement trap: celebrating lower click rates while ignoring whether users are still reporting real threats. If failure falls but reporting stays flat, the program may be improving recognition without improving response. If reporting rises but failure also rises, the program may be raising awareness of more messages but not enough judgement to resist them.

That is why resilience factor should be read alongside campaign design and message difficulty. A harder simulation can lower the factor temporarily, even when the program is healthy, so the metric is most useful when the testing method is consistent enough to support trend analysis.

Which supporting signals show the program is maturing

Beyond simulation results, a maturing program usually shows better operational follow-through. Reports should arrive faster, security teams should see cleaner report quality, and repeated mistakes should concentrate in a smaller set of scenarios. Those signals suggest the organisation is not only teaching theory, but also strengthening recognition, escalation, and memory under pressure.

It also helps to watch for false positives in reporting. If users start reporting almost everything, the program may be creating noise rather than judgement. The goal is not maximum reporting volume, but usable reporting that helps defenders triage real phishing attempts quickly and accurately.

CoPhish OAuth Token Theft via Copilot Studio is a good reminder that modern phishing often aims for more than a stolen password, so resilience has to include rapid reporting of suspicious prompts, links, and login flows, not just email awareness.

MailChimp Breach and Poland Military Breach show why a single successful social-engineering event can create outsized exposure, which is exactly why reporting speed and escalation quality matter as much as click reduction.

Risk and Threat Considerations

Awareness programs can look successful while leaving the underlying exposure intact. The main risk is false confidence: a better training score does not guarantee that users will report convincing lures, handle real login prompts safely, or slow an attacker long enough for defenders to respond.

Failure mechanism: metrics focus on attendance or simulation clicks instead of real-world reporting behaviour, so the organisation misses whether people recognise and escalate actual phishing attempts under pressure.

Impact: attackers retain a working initial access path, and the organisation may only discover the weakness after credential theft, token theft, or fraudulent payment activity has already begun.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

MITRE ATT&CK addresses the attack and risk surface, while CIS Controls v8, NIST CSF 2.0 and NIST SP 800-53 Rev 5 set the governance and control requirements practitioners need to meet.

Framework Control / Reference Relevance
CIS Controls v8 CIS-17 — Incident Response Management User reporting and triage are part of phishing response readiness.
Recommendation — Treat phishing reports as incident inputs and measure response timeliness.
NIST CSF 2.0 DE.CM-09 — Monitoring for Unauthorized Personnel, Connections, Devices, and Software Phishing awareness is validated by detection and reporting signals over time.
PR.AT-01 — Users Are Provided Awareness and Training The question asks whether awareness training is producing measurable resilience gains.
Recommendation — Track user reporting and suspicious-message monitoring as operational detection signals. Measure whether training changes behaviour, not just completion rates.
MITRE ATT&CK T1566 — Phishing The subject is resilience against phishing as an adversary access vector.
Recommendation — Map observed failures and reports to phishing techniques to improve defenses.
NIST SP 800-53 Rev 5 AT-2 — Awareness Training Training effectiveness is central to determining whether the program is improving.
Recommendation — Assess whether awareness training measurably reduces susceptibility and increases reporting.

Practitioner Guidance

What to verify: Compare simulation failure rate, report rate, and time-to-report over the same period. If only one metric improves, treat the program as incomplete and look for a measurement bias rather than assuming resilience has increased.

Decision rule: If click rates fall but reporting does not rise, prioritise reporting behaviour, triage speed, and message quality before declaring success. If reporting rises but failures do not fall, tighten scenario design and coaching around the specific lure types users still miss.

Practitioner takeaway: A phishing program is improving resilience only when it changes user action under pressure, meaning fewer successful lures, more useful reports, and a trend that holds over time.