Join our Newsletter — 33% off our NHI Course
Home FAQ Cyber Security What do security teams get wrong when they…
Cyber Security

What do security teams get wrong when they try to measure human risk with vanity metrics?

← Back to all FAQ
By NHI Mgmt Group Editorial Team Updated September 10, 2026 Domain: Cyber Security

Teams often mistake activity metrics for risk reduction, such as training completion or generic participation counts. Those numbers do not show whether behaviour changed, incidents fell, or exposure declined. A better approach is to track real-time risk scoring, targeted interventions, and measurable reductions in human-related incidents. If the metric cannot influence action, it is probably not useful.

Why Vanity Metrics Fail as Human-Risk Signals

human risk is only measurable if the metric reflects a change in exposure, judgement, or control effectiveness. Counts such as training completions, click rates, or attendance can describe programme activity, but they rarely show whether people are making safer decisions under real conditions. The danger is that teams optimise reporting comfort instead of risk reduction, which leaves the underlying exposure untouched. That is why a metric must be tied to a decision, a control, or an observable security outcome. For a broader governance lens, the NIST Cybersecurity Framework 2.0 is useful when teams need to connect measurement to outcomes rather than activity.

In practice, many security teams discover the gap only after a real incident shows that their best-looking dashboard never reflected actual behaviour change.

What Meaningful Measurement Looks Like in Practice

Useful human-risk measurement starts with the behaviour or exposure the team is trying to change, then works backwards to a metric that can support intervention. If phishing susceptibility is the concern, the important signal is not simply how many people completed awareness training, but whether risky actions declined over time and whether repeat offenders received follow-up. If unsafe sharing of credentials or data is the concern, the metric should show whether the control environment is reducing those events, not whether participation targets were met. The same principle applies to insider-risk style monitoring, privilege misuse, and policy exceptions: measure the outcome that matters, not the administrative work around it.

A practical measurement set usually combines leading and lagging indicators. Leading indicators can show whether risky conditions are improving, such as fewer repeated unsafe behaviours after intervention. Lagging indicators confirm whether incidents or near misses actually declined. Good teams also separate population-wide metrics from risk-tiered metrics so they can tell whether a programme is helping the highest-risk groups or merely lifting averages. If the data cannot drive a different intervention for a different risk segment, it is too blunt.

  • Measure behaviour change, not just participation.
  • Use metrics that support targeted action for specific user groups.
  • Check whether the metric correlates with incident reduction or exposure decline.
  • Prefer measures that can be repeated consistently over time.

This guidance breaks down when an organisation has no reliable way to observe behaviour outside simulated exercises or where downstream incidents are too rare to provide a stable signal.

Where Human-Risk Metrics Go Wrong at the Edges

Tighter measurement often increases administrative overhead, requiring organisations to balance visibility against the cost of collecting and interpreting the data. The most common edge case is using a metric that is easy to collect but too detached from actual risk, which creates confidence without control. Another is overfitting the metric to one threat type, such as phishing, and then assuming it says something general about all human risk. That kind of overreach is a design error, not a measurement success.

There is also a genuine consensus gap in the industry around whether some human-risk indicators should be used primarily for coaching, enforcement, or prioritisation. The answer depends on organisational culture and legal context, but the metric itself should still be capable of supporting a concrete decision. Vanity metrics fail most obviously when they are promoted as evidence of maturity despite not changing behaviour, reducing exposure, or informing action. The better question is whether the measure would still matter if the dashboard disappeared and the team had to defend its risk picture from incident data alone.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

NIST CSF 2.0, CIS Controls v8 and NIST AI RMF set the technical controls, while ISO/IEC 42001:2023 define the regulatory obligations.

FrameworkControl / ReferenceRelevance
NIST CSF 2.0GV.OV — OversightHuman-risk metrics belong in outcome-focused governance oversight.
GV.ME — MetricsThe question is about choosing metrics that reflect real risk change, not activity.
Recommendation — Track metrics that drive decisions and show whether risk is actually declining. Define measures that evidence control effectiveness and behaviour change.
CIS Controls v814 — Security Awareness and Skills TrainingTraining metrics are a common human-risk vanity metric and need outcome-based validation.
Recommendation — Measure training impact through behaviour and incident trends, not completion counts.
NIST AI RMFGOV-2 — AI Risk Management StrategyRisk measurement should support governance decisions, even when the subject is human behaviour.
Recommendation — Use risk metrics that inform governance actions rather than reporting activity.
ISO/IEC 42001:20239.1 — Monitoring, measurement, analysis and evaluationAI governance measurement principles apply to avoiding output-only metrics in broader risk programmes.
Recommendation — Monitor measures that demonstrate whether controls and interventions are effective.

Practitioner Guidance

What to prioritise: Start with the specific human behaviour or decision you want to reduce, then define the smallest metric that can show whether that exposure is falling. If the measure cannot distinguish between low-risk and high-risk groups, it is usually too shallow to guide action.

Decision rule: If a metric only proves participation, treat it as programme administration, not risk evidence. If it can trigger a targeted intervention, an escalation, or a control change, it is closer to a genuine risk indicator.

What practitioners underestimate: Teams often assume that more measurement equals better governance, but poorly chosen metrics can hide real exposure by making activity look like control. The useful test is whether the metric changes what the team does next.

Practitioner takeaway: Measure the behaviour you need to change, not the work you needed to schedule, because only the former tells you whether human risk is actually going down.

Deepen Your Knowledge

Sign up to our weekly newsletter — get 33% off our NHI Foundation Level Course

    NHIMG Editorial Note
    Reviewed and updated by the NHIMG editorial team on September 10, 2026.
    NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org