Security awareness teams should use an objective, repeatable scoring method based on observable phishing cues, not personal judgment. That approach helps match simulation difficulty to employee skill, reduces bias across reviewers, and produces more reliable results. When difficulty is consistent, teams can identify knowledge gaps, tailor follow-up training, and compare outcomes over time with far more confidence.
Use Objective Cues to Set Difficulty, Not Opinions
Phishing simulation difficulty should be based on observable features in the email, not a reviewer’s gut feel about the employee. A repeatable scoring rubric makes the exercise fairer, easier to defend, and far more useful for comparing results across teams, regions, and time periods. The goal is to vary challenge level by message quality, not by who happens to review it.
That means the same cues should receive the same weight every time: sender mismatch, lookalike domains, urgency language, unusual request type, attachment risk, and destination reputation. When teams anchor the score to those cues, they can explain why one simulation is basic and another is advanced, without drifting into inconsistent human judgment.
Match Difficulty to Likely Decision Pressure
The most useful difficulty scale is the one that reflects how much pressure the message creates on the recipient. Low-difficulty simulations usually contain obvious defects that most employees should notice quickly. Mid-level simulations are more realistic, but still expose common verification habits. High-difficulty simulations should mirror the kinds of messages that require careful cross-checking, not just a quick visual scan.
Difficulty also needs to account for the employee’s role and exposure. People who handle payments, approvals, external communications, or sensitive data may face more credible lures than staff with limited external contact. The point is not to make some groups “harder to fool” for its own sake, but to test them against the kinds of messages they are actually more likely to encounter.
Teams should also be careful not to let difficulty become a proxy for punishment. If the hardest simulations are used only to generate failure rates, the program can lose trust quickly. A well-designed scale distinguishes between awareness maturity and task reality, which is why the same employee may deserve a different simulation profile after a job change or a change in exposure.
Build a Calibration Process Before You Roll It Out
A scoring method only works if it is calibrated. Teams should test a sample of messages, compare scoring decisions, and look for large gaps between reviewers before launch. If one reviewer consistently treats subtle brand impersonation as “easy” and another treats it as “hard,” the simulation program will produce noisy results that are difficult to act on.
Calibration should also check whether the score predicts the outcome the program cares about. If nearly everyone fails a supposedly basic template, the scale is too aggressive. If advanced lures are routinely ignored because they are too artificial, the scale is too weak. The best programs adjust the rubric until the difficulty bands produce believable, repeatable separation in performance.
For teams building the rubric, useful background on common phishing mechanisms can be found in MITRE ATT&CK Enterprise Matrix, which helps anchor simulation design in known adversary techniques. For identity-related cues such as token theft, the NIST SP 800-63 Digital Identity Guidelines are a useful reference point for how authentication strength and phishing resistance intersect with user behavior.
Risk and Threat Considerations
If simulation difficulty is set subjectively, teams can end up measuring reviewer preference instead of employee behavior. That creates distorted metrics, weakens trend analysis, and can misclassify real risk by making one group look better or worse than it is. Over time, the bigger failure is not the test itself, but the false confidence created by inconsistent scoring.
Failure mechanism: Inconsistent difficulty scoring changes the baseline from one reviewer or department to another, so results are no longer comparable and the training response becomes misaligned.
Impact: Teams may overtrain low-risk groups, undertrain high-risk groups, and miss the patterns that would have justified targeted follow-up or control improvements.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
MITRE ATT&CK addresses the attack and risk surface, while NIST SP 800-63, NIST CSF 2.0, NIST SP 800-53 Rev 5 and CIS Controls v8 set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| MITRE ATT&CK | T1566 — Phishing | Phishing simulation difficulty is anchored to phishing techniques and lures. |
| Recommendation — Map simulation cues to phishing techniques and tune scenarios to the adversary patterns you want to test. | ||
| NIST SP 800-63 | Digital Identity Guidelines | Phishing-resistant authentication informs which users need harder simulations. |
| Recommendation — Use phishing-resistant authentication guidance to align simulation difficulty with authentication risk. | ||
| NIST CSF 2.0 | PR.AT-01 — Awareness and Training | The topic is directly about designing effective security awareness training exercises. |
| Recommendation — Calibrate awareness exercises so they measure behavior consistently across employee groups. | ||
| NIST SP 800-53 Rev 5 | AT-2 — Awareness Training | Phishing simulations are a core awareness-training activity under formal security controls. |
| Recommendation — Align simulation difficulty with the training outcome you want to validate. | ||
| CIS Controls v8 | CIS-14 — Security Awareness and Skills Training | The question is about operationalizing awareness training for phishing resistance. |
| Recommendation — Standardize phishing simulations as part of a repeatable awareness and skills program. | ||
Practitioner Guidance
What to verify: Before trusting a difficulty tier, verify that at least two reviewers score the same sample messages within a narrow range and that the rubric distinguishes obvious defects from subtle, realistic lures. If the score cannot survive reviewer comparison, it is not ready for operational use.
Decision rule: If the simulation is meant to measure awareness behavior, keep the scoring tied to observable cues and role exposure; if it is meant to test a specific control, such as reporting or verification behavior, tune the difficulty to that control rather than to generic “hardness.”
Practitioner takeaway: The best phishing difficulty model is the one that is repeatable enough to trust, specific enough to explain, and stable enough to support trend data without turning employee scoring into a subjective exercise.
Related resources from NHI Mgmt Group
- How should security teams choose identity verification controls for different risk levels?
- How should security teams choose between DV, OV, and EV certificates for different website risk levels?
- How should security teams implement human risk management in environments where employees have different access levels and threat exposure?
- How should security teams choose between e-signatures and digital signatures for different document risk levels?