By NHI Mgmt Group Editorial TeamDomain: Cyber SecuritySource: SawmillsPublished September 11, 2025

TL;DR: Only 13% of collected telemetry is used, while 84% of companies consume less than a quarter of what they ingest and average annual observability spend reaches $905,000, according to Sawmills’ 2025 State of Observability and Telemetry Report. The governance problem is no longer visibility alone, but how to control telemetry volume, quality, and cost without blinding incident response.


At a glance

What this is: The report finds that observability pipelines are overloaded with redundant telemetry, and most collected data is never used.

Why it matters: This matters because teams responsible for identity, platform, and security operations depend on high-signal telemetry to detect misuse, investigate access events, and keep response fast.

By the numbers:

👉 Read Sawmills' 2025 State of Observability and Telemetry Report


Context

Observability data waste is what happens when teams collect logs, traces, and metrics faster than they can govern, filter, and use them. The result is not better visibility, but more noise, higher cost, and slower troubleshooting across engineering and security operations.

For identity and security practitioners, the issue is not abstract. Telemetry sprawl can obscure access misuse, delay investigation of service account abuse, and weaken the signal needed for IAM, PAM, and SOC workflows. The pattern described here is increasingly typical in large enterprises rather than an edge case.


Key questions

Q: How should teams reduce observability costs without losing useful telemetry?

A: Start at the pipeline, not the backend. Filter obvious noise, remove high-cardinality fields, and tail-sample traces so routine traffic is reduced after full context is available. That preserves error and latency signal while cutting ingest, storage, and query costs before they become locked into the observability bill.

Q: Why do observability endpoints create governance issues for platform teams?

A: Because metrics endpoints reveal service behaviour, error patterns, and deployment structure, which can help both operators and attackers map the environment. Once access is broad, the organisation has created a data surface that needs ownership, policy, and review, not just engineering convenience.

Q: What do security and engineering teams get wrong about collecting more telemetry?

A: They often assume that more data automatically means better detection. In reality, unfiltered telemetry can bury the evidence that matters and increase the time needed to resolve incidents. If teams cannot answer which data is used, by whom, and for what purpose, they are collecting noise, not observability.

Q: How do organisations know if telemetry governance is working?

A: Look for fewer unnecessary ingestions, higher-value events reaching the SIEM, and clearer ownership of collection rules and routing logic. A working programme can explain why data is collected, who can change it, and how enrichment improves both cost and detection outcomes.


Technical breakdown

Why telemetry volume breaks signal quality

Modern observability pipelines ingest logs, traces, and metrics from applications, infrastructure, and cloud services at scale. The problem is not collection itself, but ungoverned volume, duplicate events, and high-cardinality data that expands faster than analysts can triage it. When the pipeline lacks filtering, enrichment, and retention policy, the result is signal dilution. Teams see more dashboards and alerts, but less operational clarity. In practice, this means the most relevant anomaly can be buried under repetitive low-value telemetry.

Practical implication: establish filtering and retention rules before adding more collection sources.

How tool sprawl turns observability into a governance problem

Running multiple observability platforms often creates overlapping ingestion, duplicated storage, inconsistent alerting, and fragmented ownership. Once data is split across tools, it becomes harder to correlate events across runtime, cloud, application, and identity layers. That fragmentation is a governance issue, not just a tooling preference, because it affects who can see what, how long data is kept, and whether investigations can reconstruct an incident end to end. The cost increase is a symptom of control drift.

Practical implication: rationalise overlapping platforms and assign a single owner for telemetry governance.

Where AI can help telemetry operations, and where it cannot

AI copilots and AI agents can reduce manual triage by clustering anomalies, summarising incidents, and prioritising likely root causes. But they do not fix poor data quality, inconsistent schemas, or weak access boundaries around telemetry stores. If the underlying observability estate is noisy, AI will accelerate the noise rather than remove it. The real value comes when AI is constrained by governed data pipelines, clear policy, and human review for high-impact actions.

Practical implication: use AI to augment triage only after telemetry quality and access controls are in place.


NHI Mgmt Group analysis

Observability sprawl is becoming a control-plane problem, not a tooling problem. Once multiple platforms ingest overlapping telemetry, organisations lose consistency in retention, correlation, and escalation paths. That weakens both operational resilience and security investigation quality, especially when access events and infrastructure anomalies need to be joined quickly. The practical conclusion is that telemetry governance now belongs in the same conversation as platform governance.

High-volume telemetry without policy creates a signal-to-noise debt. The cost is not just financial, although the report’s spend figures make that obvious. The deeper issue is that analysts and engineers spend more time sifting than deciding, which pushes incident resolution into a slower, less reliable mode. The named concept here is signal-to-noise debt: the longer teams defer telemetry discipline, the more expensive every investigation becomes.

Identity and access investigations depend on clean telemetry more than most teams acknowledge. Access misuse, service account abuse, and privilege escalation often appear first as subtle log patterns rather than obvious incidents. If the telemetry estate is bloated or fragmented, those patterns disappear into background noise. For IAM and PAM teams, observability hygiene is part of detection readiness, not a separate engineering concern.

AI-assisted observability will only help if the data pipeline is already governed. The report’s interest in AI agents and copilots reflects where the market is heading, but AI is an amplifier, not a substitute for policy. Teams that automate triage without fixing data quality will simply automate confusion faster. Practitioners should treat AI as a layer on top of observability governance, not a replacement for it.

What this signals

Observability programmes are moving from collection-heavy designs toward governance-heavy designs. The practical signal for practitioners is that telemetry policy, schema discipline, and platform rationalisation will matter more than adding another dashboard or AI layer.

Signal-to-noise debt: teams that keep expanding ingestion without pruning low-value sources will make every later incident harder to investigate. The operational consequence is slower containment, weaker correlation, and more time spent on search than on decision-making.

Identity telemetry should be treated as a first-class investigative dataset. Service account activity, authentication anomalies, and privileged sessions only remain useful if the pipeline preserves them with clear retention, access, and normalisation rules.


For practitioners

  • Measure telemetry utility, not just ingestion volume Track the percentage of collected telemetry that is actually queried, correlated, or used in incident response. Separate high-value sources from redundant ones and use that data to prune low-yield pipelines.
  • Consolidate overlapping observability platforms Inventory where logs, traces, and metrics are duplicated across platforms, then designate one system of record for each telemetry class. This reduces fragmentation and clarifies ownership for retention, alerting, and access.
  • Protect identity-related telemetry as an investigation asset Ensure logs tied to service accounts, privileged access, and authentication events are retained, normalised, and searchable long enough to support root-cause analysis and access abuse investigations.
  • Introduce policy before adding AI to observability workflows Use AI to summarise and prioritise only after schema quality, access controls, and filtering rules are in place. Otherwise the model will surface repetitive noise rather than actionable anomalies.

Key takeaways

  • Observability costs are rising because organisations keep ingesting more telemetry than they can actually use.
  • The report’s figures show that telemetry waste is now a governance issue as much as a budget issue.
  • Teams need to control data volume, platform sprawl, and AI usage together if they want faster, more reliable incident response.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

NIST CSF 2.0, NIST SP 800-53 Rev 5, CIS Controls v8 and NIST IR 8596 set the technical controls, while ISO/IEC 27001:2022 define the regulatory obligations.

FrameworkControl / ReferenceRelevance
NIST CSF 2.0DE.CM-1Telemetry quality and monitoring coverage are central to continuous detection and response.
NIST SP 800-53 Rev 5AU-6Telemetry noise and retention discipline map directly to audit review and analysis controls.
CIS Controls v8CIS-8 , Audit Log ManagementThe report’s problem is audit-log overload, fragmentation, and poor utility.
ISO/IEC 27001:2022A.8.15Logging and monitoring controls are directly implicated by telemetry governance and retention quality.
NIST IR 8596AI copilots for SRE and telemetry optimisation connect to AI system governance and operational use.

Use CIS-8 to standardise logging, retention, and review so telemetry supports investigations instead of obscuring them.


Key terms

  • Telemetry-driven governance: Telemetry-driven governance is a control approach that relies on runtime signals rather than periodic paperwork. For AI, that means watching drift, leakage, prompt anomalies, and other live indicators so governance decisions reflect current system behaviour instead of stale review findings.
  • Signal-to-noise debt: Signal-to-noise debt is the accumulated cost of collecting more operational data than teams can efficiently interpret. As it grows, the useful signals become harder to isolate, incident response slows, and the organisation pays more for less decision value.
  • Security Observability Sprawl: Security observability sprawl is the condition where useful evidence is spread across too many disconnected tools, formats, and workflows. It raises operational cost because teams must stitch together context before they can decide whether an event is real, relevant, or escalating.
  • Mean Time To Resolution: Mean time to resolution is the average time it takes a supplier to fix an issue from the moment it is reported or detected. It is a useful service metric because it shows not just whether something broke, but how quickly the vendor can restore reliable operation.

What's in the full report

Sawmills' full report covers the operational detail this post intentionally leaves for the source:

  • Survey methodology and respondent breakdown across US and EU senior DevOps and engineering leaders
  • Platform-by-platform observability stack patterns showing where tool sprawl is most common
  • Detailed spending breakdowns behind the average $905,000 annual observability bill
  • AI adoption findings on copilots and agents for telemetry optimisation and incident response

👉 The full Sawmills report includes the survey data, platform mix, and AI adoption findings behind the cost and telemetry waste trends.

Deepen your knowledge

The NHI Foundation Level course, the industry's only accredited NHI security programme, covers NHI governance, machine identity security, and secrets management. It helps practitioners connect identity controls to the broader security programmes they already run.
NHIMG Editorial Note
Published by the NHIMG editorial team on August 19, 2026.
NHI Mgmt Group — the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org