Annual training can raise awareness, but it rarely proves whether people will act safely under pressure. Social engineering tests reveal whether training changes behaviour, whether employees report suspicious activity, and where processes break down in real conditions. That makes testing a measurement tool for human risk, not just a compliance exercise.
Why Social Engineering Tests Still Matter After Awareness Training
Annual awareness training can improve baseline knowledge, but it does not prove that people will recognise pressure tactics in the moment or follow the right escalation path. social engineering tests measure behaviour, not attendance. That matters because attackers exploit urgency, authority, curiosity, and routine, which are difficult to validate in a classroom setting. NIST’s identity guidance emphasises that authentication and assurance depend on real-world verification, not assumptions formed in training alone, as reflected in the NIST SP 800-63 Digital Identity Guidelines.
For security teams, the test is also a process check. A convincing phishing email, phone call, or text can reveal whether employees pause, verify, and report, or whether they bypass controls under pressure. NHIMG has documented how social engineering has repeatedly been the entry point in major incidents such as MGM Resorts Breach 2023 and the Storm-2949 Azure Breach, where human trust and workflow gaps were central to compromise. In practice, many security teams discover those gaps only after an attacker has already used them, rather than through intentional testing.
How Social Engineering Testing Measures Behaviour, Not Just Knowledge
Effective testing works because it mirrors attacker tradecraft closely enough to expose decision points without turning the exercise into a simple memory quiz. A useful program tests multiple channels, such as email, voice, SMS, and help desk interactions, then measures whether people report, challenge, or comply. It should also assess whether downstream controls, including ticketing, callback verification, and manager approval, actually slow unsafe actions. NIST’s control catalog supports this kind of operational verification through monitoring and incident response expectations in NIST SP 800-53 Rev 5 Security and Privacy Controls.
Current best practice is to score more than click rates. Mature programs track reporting rates, time to report, identity verification failures, and repeat exposure by team or function. That creates a behavioural baseline that annual training cannot provide on its own. It also helps teams distinguish between users who need refresher training and workflows that are structurally unsafe. NHIMG incident research on Co-op Group DragonForce Breach and Uber Breach shows why this matters: once an attacker can manipulate a trusted workflow, a single human action can unlock broad access.
- Use realistic lures that match current threats, not generic templates.
- Measure reporting and escalation, not only failure or success.
- Test the help desk, executives, and contractors as separate risk paths.
- Review whether written procedures are actually usable under pressure.
- Feed results into process fixes, not just awareness reminders.
These controls tend to break down in highly decentralised environments where local exceptions, informal approvals, and inconsistent identity verification make the “right” response unclear in the moment.
Where the Standard Answer Breaks Down in Real Organisations
Tighter testing often increases administrative overhead and can create alert fatigue, so organisations have to balance behavioural insight against employee trust and operational cost. There is also a genuine tradeoff between realism and ethics: current guidance suggests avoiding exercises that shame staff or collect unnecessary personal data, because those practices can reduce reporting quality over time. The goal is evidence, not embarrassment.
One common edge case is mature awareness programs with weak process design. In those environments, employees may know the policy but still fail because the workflow requires speed, repeated approvals, or ambiguous ownership. Another edge case is fast-moving threat activity. NHIMG’s analysis of the DeepSeek breach and vendor research from The State of Secrets in AppSec show how quickly exposed credentials and weak handling practices can turn into real loss, which is why social engineering tests should be paired with fast reporting and containment workflows.
There is no universal standard for how often to test or how aggressive the scenarios should be. The best programs align test design to the organisation’s actual attack surface, then use the results to improve verification, escalation, and access control. In other words, training teaches the rule, but testing proves whether the rule survives contact with the real environment.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Non-Human Identity Top 10 address the attack and risk surface, while NIST CSF 2.0, NIST SP 800-63 and NIST AI RMF set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| NIST CSF 2.0 | PR.AT-1 | Awareness and training are relevant, but testing checks whether training changes behaviour. |
| NIST SP 800-63 | Identity assurance depends on real verification, which social engineering tries to bypass. | |
| OWASP Non-Human Identity Top 10 | NHI-05 | Human-approved workflows often expose credentials and access that attackers target. |
| NIST AI RMF | GOVERN | Testing supports accountability for human and process risk in security operations. |
Assign ownership for test design, remediation, and behavioural metrics under AI RMF governance.
Related resources from NHI Mgmt Group
- Why do phishing and social engineering remain so effective against Web3 organisations?
- How do security and fraud teams measure whether awareness training is actually reducing social engineering risk?
- Why do non-human identities create compliance risk even when policies exist?
- Why do leaked secrets remain such a persistent NHI risk?
Deepen Your Knowledge
Reviewed and updated by the NHIMG editorial team on August 23, 2026.
NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org