Generic simulations teach employees to recognize the test, not the threat. When the same templates are reused, the program loses realism, underestimates actual risk, and fails to measure response to modern tactics like business email compromise, credential harvesting, QR code phishing, and SMS phishing. A useful program must evolve with current threat intelligence and mixed delivery channels.
Why This Matters for Security Teams
Generic phishing tests create a false sense of maturity. If employees can spot the simulation because of reused wording, obvious sender patterns, or predictable landing pages, the program measures familiarity with the exercise rather than resilience against real adversaries. That matters because modern phishing is not a single tactic: it blends credential theft, conversation hijacking, QR codes, SMS, collaboration platforms, and identity abuse. The security objective is not to “catch” people out, but to understand whether the organisation can recognise and interrupt a live attack path.
This is where control design matters. The NIST Cybersecurity Framework 2.0 emphasises governance, awareness, and continuous improvement, which is exactly what predictable testing often lacks. If the exercise never varies by channel, urgency, business context, or attacker objective, the lessons become procedural rather than behavioural. Security teams then overestimate reporting quality, under-measure clickthrough risk, and miss whether users escalate suspicious requests through the right channel. In practice, many security teams discover this only after a real business email compromise has already moved from inbox to payment request or account takeover.
How It Works in Practice
A credible phishing programme should test judgment, not memory. That means varying the lure, the delivery method, the target role, and the expected action. A finance user should not see the same template as a help desk analyst or an executive assistant, because real attackers tailor pretexts to privilege and workflow. Current guidance suggests aligning simulations with actual threat patterns seen in the organisation’s sector, including credential harvesting, invoice fraud, OAuth consent abuse, SMS-based prompts, and QR code redirection.
Practically, that requires more than periodic email blasts. Effective programmes usually combine:
- Threat-informed templates based on recent intelligence and internal incident trends.
- Multi-channel delivery, including email, SMS, collaboration tools, and voice when relevant.
- Behavioural outcomes, such as reporting speed, escalation quality, and identity verification steps taken.
- Controls around realism, so simulations do not create unsafe links, brand misuse, or unnecessary data exposure.
Teams should also connect testing to response workflows. If an employee reports a suspicious message, the security function should be able to triage it in SIEM, validate related indicators, and contain related accounts or sessions quickly. That makes the exercise useful for both awareness and operational readiness. Where identity governance is involved, the question is whether users are trained to verify high-risk requests through trusted channels before approving access, payments, or credential resets. The MITRE ATT&CK framework is useful here because it helps map phishing to downstream techniques such as valid account use, token theft, and initial access chains. These controls tend to break down in very large, decentralised organisations because template governance, localisation, and business-unit exceptions make consistency hard to maintain.
Common Variations and Edge Cases
Tighter phishing realism often increases programme complexity, requiring organisations to balance fidelity against employee trust and administrative overhead. There is no universal standard for how deceptive a simulation should be, and best practice is evolving around ethics, consent, and regional labour requirements. The right answer usually depends on the organisation’s risk profile, legal environment, and maturity of the broader awareness programme.
Some edge cases need special handling. High-trust environments such as executive support, procurement, and payroll often warrant more sophisticated scenarios because the attacker payoff is higher. Remote-first organisations may need more emphasis on collaboration-platform lures and mobile delivery. Regulated sectors may need stronger controls on data use, logging, and approval chains. The OWASP phishing guidance is useful for shaping tests that reflect realistic web and credential risks, while the CISA phishing resources help anchor the exercise in current attacker behaviour.
For identity-heavy workflows, the key edge case is whether a suspicious message leads to account recovery, MFA reset, or privileged access approval. That is where generic testing fails most visibly: users may recognise a fake email, but still approve a malicious prompt or hand over a one-time code. In that environment, the programme should test decision points, not just inbox recognition.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
MITRE ATT&CK and OWASP Agentic AI Top 10 address the attack and risk surface, while NIST CSF 2.0, NIST AI RMF and NIST AI 600-1 set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| NIST CSF 2.0 | PR.AT-01 | Awareness and training should build real phishing resilience, not template recognition. |
| MITRE ATT&CK | T1566 | Phishing is the core initial-access technique these tests should emulate. |
| OWASP Agentic AI Top 10 | Predictable social engineering can also target AI assistants and workflow agents. | |
| NIST AI RMF | Realistic testing should support governance, measurement, and continuous improvement. | |
| NIST AI 600-1 | If AI tools draft or send lures, the content pipeline itself needs validation. |
Use awareness controls to measure and improve user response to realistic attack patterns.
Related resources from NHI Mgmt Group
- What breaks when approval rules are too generic for different identity types?
- What breaks when identity and access policies are too generic for frontline workflows?
- Why do credential phishing simulations matter more than generic awareness tests?
- What breaks when security tools are too generic for the code they scan?
Deepen Your Knowledge
Reviewed and updated by the NHIMG editorial team on August 24, 2026.
NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org