Join our Newsletter — 33% off our NHI Course
Home FAQ Cyber Security Why can a model with better average results…
Cyber Security

Why can a model with better average results still be the wrong choice for a security workflow?

← Back to all FAQ
By NHI Mgmt Group Editorial Team Updated September 20, 2026 Domain: Cyber Security

A model with a better average score can still be wrong if its marginal improvement is weaker at the operating point that matters. What matters is the incremental gain from adding effort, budget, or threshold changes in context. In practice, security teams should compare the slope of the trade-off curve, not just the headline averages.

Why the average can mislead in security decisions

A security workflow rarely cares about the average case in isolation. What matters is whether a model improves the specific operating point you use, such as a high-recall threshold for triage or a low-false-positive threshold for blocking. A model can look better overall and still deliver less useful gains where the workflow actually spends time, budget, or risk tolerance.

That is why practitioners should treat summary metrics as screening data, not as the decision itself. In a real workflow, the question is whether the model gives you a better exchange of effort for risk reduction at the point where humans or controls intervene. If the gain flattens early, the headline average can hide the fact that the model is expensive to tune for little practical benefit.

When evaluating models used for detection, classification, or prioritisation, it helps to compare the shape of the trade-off curve rather than a single score. A small improvement in average accuracy may be less valuable than a larger gain at the threshold that determines alert volume, escalation rate, or analyst workload.

Where the operating point changes the answer

The right model depends on the cost of errors in context. In a workflow that must catch rare but high-impact events, the most useful model is often the one that improves the left or right edge of the curve, not the one that wins on aggregate. If changing the threshold or adding review effort produces only a weak marginal gain, you are paying for complexity without meaningfully improving security outcomes.

This is especially important when the model is used as one step in a larger control chain. A slightly better average result can still be the wrong choice if it increases manual review, slows containment, or shifts noise into the next stage of the process. The practical decision is not “which model scores highest”, but “which model gives the best incremental return at the decision boundary we actually use”.

For teams comparing candidates, the most useful evidence is often a side-by-side view of precision, recall, and workload at the intended threshold. That makes the trade-off visible, which is more useful than a single headline metric that averages away the exact region where the workflow is sensitive.

For background on how average performance can be undermined by the practical distribution of security problems, NHI Mgmt Group’s Ultimate Guide to NHIs research and survey results shows how scale and visibility gaps change the value of a control in practice, and the 52 NHI breaches analysis is useful when you want to understand how control failure can matter more than a nominal average gain.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

NIST CSF 2.0, CIS Controls v8 and NIST AI RMF set the governance and control requirements practitioners need to meet.

FrameworkControl / ReferenceRelevance
NIST CSF 2.0ID.IM-01 — Improvements Are Identified and MadeModel selection should improve security decisions at the actual operating point.
Recommendation — Assess whether the model improves decisions at the threshold your workflow really uses.
CIS Controls v88 — Audit Log ManagementDecision quality in security workflows depends on measurable, reviewable operational evidence.
Recommendation — Measure workflow impact with logged outcomes, not just headline model scores.
NIST AI RMFMAP-A — MapChoosing a model by operational trade-off fits AI risk assessment and context mapping.
Recommendation — Map the model to the security task, loss profile, and decision boundary before comparing scores.

Practitioner Guidance

What to verify: Compare models at the exact threshold, review queue size, or decision rule the workflow will actually use. If two models differ by only a small average score but one materially reduces analyst load or false positives at the chosen operating point, that is the stronger security choice.

Decision rule: Prefer the model with the better marginal gain where the workflow spends effort, not the model with the best overall average. If the curve is flatter near your operating point, treat the model as weaker even if the benchmark score looks superior.

Common mistake: Teams often optimise for benchmark reporting and then discover that the “better” model is harder to tune, slower to run, or noisier in production. The output that matters is the security decision under real constraints, not the ranking on a generic test set.

Practitioner takeaway: In security, a model is only as good as the improvement it delivers at the decision boundary that matters, so evaluate incremental value, not just aggregate performance.

Deepen Your Knowledge

Sign up to our weekly newsletter — get 33% off our NHI Foundation Level Course

    NHIMG Editorial Note
    Reviewed and updated by the NHIMG editorial team on September 20, 2026.
    NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org