Collecting more data without clear use cases usually creates a data swamp. Analytics lag behind ingestion, most of the data remains unused, and low-quality fields make detections fragile. That increases manual tuning, raises false positives, and burns analyst time. In practice, volume alone does not improve security outcomes unless the data is structured, relevant, and operationally usable.
Why More Telemetry Can Degrade Detection Quality
Security teams often assume that better detection comes from collecting everything, but detection systems are constrained by ingest cost, parser quality, normalization, and analyst capacity. Once a pipeline is flooded with low-value or inconsistent events, signal gets buried under noise and detections become harder to trust. The practical problem is not lack of data, but lack of disciplined selection and operational fit. The NIST Cybersecurity Framework 2.0 is useful here because it emphasises outcomes, not raw volume, which is the right lens for judging whether telemetry is improving security.
In practice, many security teams discover this only after they have expanded collection faster than they have improved normalisation, enrichment, and alert triage.
How Overcollection Breaks the Detection Pipeline
Detection quality depends on the full path from data source to decision, not just on how many events arrive. If a team adds logs without validating schema consistency, field completeness, and business relevance, it increases the number of cases the pipeline must handle while reducing the fraction that can be interpreted reliably. That creates a compound failure: analytics spend more time sorting irrelevant records, correlation rules become brittle, and meaningful anomalies are harder to distinguish from routine system churn.
More data also raises the maintenance burden. Every new source needs parsing, mapping, retention rules, storage planning, and content tuning. If those tasks are not resourced, teams end up with partial visibility that looks broad on paper but performs poorly in investigation. A log source that is rarely queried may still consume budget and operational attention, while a smaller set of high-fidelity sources can support faster and more accurate detections.
- Relevant data improves detections when it is consistent, timely, and mapped to a clear use case.
- Irrelevant data worsens detections when it increases false positives or masks the patterns analysts need.
- Low-quality fields undermine correlation because detections become dependent on missing or unreliable attributes.
- Excess collection can delay alerting when storage, enrichment, or routing becomes a bottleneck.
For that reason, mature detection engineering usually starts with the question of which behaviours must be seen, then works backward to the minimum telemetry needed to see them well. Teams that reverse that order tend to build broad collection estates that are expensive to run and still insufficient for reliable detection. The guidance breaks down when a specific compliance or forensic requirement truly demands high-volume retention, because then the goal is evidentiary coverage rather than day-to-day detection quality.
When Extra Data Helps and When It Just Adds Noise
Tighter collection often improves visibility only when it closes a known gap, which means organisations must balance improved context against the overhead of maintaining it. The key tradeoff is between breadth and usability: a smaller, well-governed set of sources usually supports stronger detection than a larger pile of weakly managed telemetry. The correct answer is not “collect less” in all cases, but “collect with a decision in mind.”
There is still a role for broader data in investigations, threat hunting, and long-horizon retention, but those use cases should not be confused with operational detection. If a source cannot be tied to a named analytic, a response workflow, or a regulatory purpose, it is usually a candidate for reduction, tiering, or delayed ingestion rather than immediate promotion into the primary detection stack. Where teams disagree, the useful test is whether the data changes a decision, not whether it merely increases the size of the dataset.
Practitioners also underestimate how quickly duplicate or overlapping sources create confidence problems. If several tools report the same event with different timestamps, actor labels, or object names, triage becomes slower instead of faster. The result is not richer detection, but more reconciliation work before anyone can trust the alert.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
MITRE ATT&CK address the attack and risk surface, while NIST CSF 2.0 and CIS Controls v8 set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| NIST CSF 2.0 | DE.CM-01 — Monitoring for Anomalies and Events | Detection quality depends on meaningful monitoring outputs, not raw telemetry volume. |
| ID.RA-01 — Risk Identification | Overcollection creates operational risk when data is not tied to a clear security use case. | |
| RS.AN-01 — Analysis of Events and Alerts | Alert triage degrades when too much low-value data reaches analysts. | |
| Recommendation — Prioritise telemetry that improves anomaly detection and response decisions. Tie each source to a documented risk or detection purpose before expanding collection. Measure whether added data reduces or increases analyst effort per actionable alert. | ||
| CIS Controls v8 | 8 — Audit Log Management | The issue is log usefulness, quality, and manageability, not mere collection. |
| Recommendation — Curate logging so collected events remain usable for investigation and detection. | ||
| MITRE ATT&CK | T1119 — Automated Collection | Excess collection can expand observable surface while overwhelming defenders with noise. |
| Recommendation — Map collection coverage to defender value and watch for telemetry that adds little analytic utility. | ||
Practitioner Guidance
What to prioritise: Start with the detections you actually need to run, then inventory the minimum telemetry required to support them. If a source does not improve an existing use case, support a response workflow, or close a known visibility gap, it should not be treated as core detection data.
What to verify: Check whether the pipeline can parse, normalise, and enrich the data at the same pace it is ingested. A source that arrives faster than it can be made queryable is operational debt, not improved visibility.
Common mistake: Treating ingestion volume as a proxy for maturity. In reality, detection quality is usually limited by schema quality, correlation design, and analyst workload long before it is limited by the amount of raw telemetry.
Practitioner takeaway: The most effective detection programmes curate data to support decisions; they do not assume that more events automatically produce better security judgement.
Related resources from NHI Mgmt Group
- Why do access bottlenecks often make security outcomes worse instead of better?
- Why do retries sometimes make outages worse instead of better?
- How should security teams use identity data for threat detection instead of just compliance reporting?
- Why do periodic password changes often make security worse?
Deepen Your Knowledge
Reviewed and updated by the NHIMG editorial team on September 9, 2026.
NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org