Join our Newsletter — 33% off our NHI Course
Home Glossary Cyber Security False Discovery Rate
Cyber Security

False Discovery Rate

← Back to Glossary
By NHI Mgmt Group Updated September 20, 2026 Domain: Cyber Security

False discovery rate is the proportion of predicted results that are wrong. It is a practical way to estimate how many false positives a model will generate at a given alert volume. For security teams, FDR is often more operationally useful than precision when comparing expected alert burden.

False discovery rate, or FDR, describes the share of predicted positives that turn out to be wrong. That makes it especially useful when the question is not just “how accurate is the model?” but “how much wasted analyst effort will this alert stream create at this operating point?”

In security operations, FDR helps compare models or rules at the same alert volume. A system can look strong on paper yet still generate too many unnecessary investigations if its false discovery rate is high, so FDR is often a better operational lens than a single abstract quality score.

FDR is closely related to precision, but the operational interpretation is different. Precision tells you how often flagged items are right; FDR tells you how often they are wrong. For teams that must size queues, triage capacity, and escalation thresholds, that distinction matters more than a purely statistical one.

Why FDR matters in security alerting

FDR is most valuable when alerts are expensive to inspect. In detection engineering, threat hunting, and fraud or abuse monitoring, a lower FDR usually means less alert fatigue, better analyst throughput, and fewer false escalations that distract from real incidents.

It also helps when comparing controls that produce different alert distributions. Two detectors may have similar recall, yet one may bury a team in noise while the other preserves operational attention for the cases most likely to be real. FDR makes that burden visible.

For this reason, FDR is often a practical bridge between model evaluation and operations. It converts an abstract classification result into a workload question: among the alerts the team will actually see, how many are expected to be wrong?

How to interpret FDR correctly

FDR should always be read alongside the threshold, dataset, and base rates that produced it. Changing the decision threshold can improve or worsen FDR quickly, so a good number at one operating point may not hold at another. In security, where positive events are often rare, base-rate effects can make the alert stream look better or worse than intuition suggests.

It is also important to remember that FDR is a population-level measure, not a guarantee about any single alert. A low FDR still means some alerts are false, and a higher FDR may be acceptable if the detections catch high-value events that justify the investigation cost.

When teams use FDR well, they treat it as part of alert governance rather than a stand-alone score. The metric should support decisions about thresholds, suppression, tuning, and analyst coverage, not replace those decisions.

What good FDR practice looks like

Choose FDR when the business problem is alert burden, not just statistical correctness. It works best when paired with volume, precision, recall, and downstream investigation cost, because no single metric captures both detection value and operational load.

Use the metric at the same decision threshold you intend to run in production. Comparing two detectors at different thresholds can be misleading if one looks better only because it is less aggressive. The useful question is how much wrong work the team will inherit at the alert level it will actually operate.

For security teams, the key judgment is whether a change reduces noise without hiding important events. That is why FDR belongs in tuning conversations, triage design, and control selection, especially when analyst time is the limiting resource.

Risk and Threat Considerations

High false discovery rate creates a concrete operational risk: analysts spend more time on noise, real attacks wait longer in queue, and repeated false alerts can erode trust in detections. Over time, teams may suppress, ignore, or under-prioritise signals that are actually valuable.

Failure mechanism: poor thresholding, weak feature quality, or an imbalanced event distribution produces too many incorrect positives at the chosen operating point, which inflates alert volume without improving true detection value.

Impact: alert fatigue, slower response, wasted investigation effort, and degraded confidence in the detection program can follow, especially when the same noisy control is used across many assets or identities.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

NIST CSF 2.0, CIS Controls v8 and NIST AI RMF set the governance and control requirements practitioners need to meet.

FrameworkControl / ReferenceRelevance
NIST CSF 2.0DE.CM — Continuous MonitoringFDR informs how well alerts support ongoing detection monitoring.
DE.AE — Anomalies and EventsFDR measures how often event-based detections are wrong.
Recommendation — Tune alert thresholds and detection monitoring to reduce false positives at the operating point you actually run. Validate event-detection rules against real event outcomes so alert quality matches operational burden.
CIS Controls v88 — Audit Log ManagementLog-derived detections often generate alerts whose false discovery rate affects analyst workload.
17 — Incident Response ManagementHigh FDR directly affects triage and investigation efficiency during incident response.
Recommendation — Review log-based alerting rules to reduce noise before they consume analyst capacity. Use alert-quality metrics to prioritise detections that support faster, more reliable incident triage.
NIST AI RMFMAP — Measure AI SystemsFDR is a measurement lens for evaluating model outputs at a chosen threshold.
MANAGE — Manage AI RisksOperational use of FDR supports governance decisions about acceptable model error and burden.
Recommendation — Measure model outputs at the production threshold to understand false-positive burden before deployment. Set governance thresholds for acceptable false-discovery burden before relying on a model operationally.

Practitioner Guidance

What to watch for: treat FDR as a tuning and governance metric, not a score to optimise in isolation. The useful operating question is whether the alert stream remains actionable at the volume your team can actually investigate.

Practitioner takeaway: if a detector’s FDR is acceptable only in a lab but not in production, it is not ready for operational use, even if its headline performance looks strong.

Deepen Your Knowledge

Sign up to our weekly newsletter — get 33% off our NHI Foundation Level Course

    NHIMG Editorial Note
    Reviewed and updated by the NHIMG editorial team on September 20, 2026.
    NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org