Common signs include rising storage and processing bills, growing dependence on specialised staff, duplicated tooling for the same telemetry types, and slower responses to changes in the environment. Another warning sign is heavy agent overhead across many systems. When teams spend more time maintaining the pipeline than using the data, the observability model is no longer efficient.
Why This Matters for Security Teams
An observability platform can quietly shift from a strategic control to an operational burden when its cost curve grows faster than the value of the telemetry it produces. For security teams, that matters because visibility is not free: storage, ingestion, parsing, retention, analyst time, and pipeline maintenance all compete with budget for detection engineering and response. A platform that is too expensive at scale often forces painful trade-offs, such as reducing retention, narrowing coverage, or delaying onboarding of new assets.
That pressure is especially important in environments where logging supports incident response, compliance, and threat hunting at the same time. If cost controls are applied too late, teams may find that the data they cut was the data needed to investigate a breach or prove control effectiveness. Current guidance suggests treating telemetry as a governed security capability rather than an unlimited utility. NIST SP 800-53 Rev 5 Security and Privacy Controls is useful here because it anchors logging, monitoring, and retention decisions in control intent rather than convenience.
In practice, many security teams notice the problem only after a budget review or incident review exposes how much effort is being spent to keep the pipeline alive rather than to improve detection quality.
How It Works in Practice
Cost problems usually emerge across four layers: data volume, pipeline complexity, human effort, and duplicated capability. High-cardinality logs, verbose traces, and long retention periods can drive ingestion and storage costs upward. At the same time, schema changes, noisy agents, and repeated normalization work increase engineering overhead. The platform may still be “working,” but the economics no longer match the scale of the environment.
A practical assessment starts with separating signal from utility. Teams should ask which data types directly support incident response, which exist mainly for troubleshooting, and which are duplicated elsewhere. They should also measure whether the platform’s cost is concentrated in a few high-value sources or spread across many low-value ones. If the expensive parts are also the least used, the platform is drifting away from security value.
- Review ingestion by source, not just total spend, so noisy systems are visible.
- Check whether retention periods are aligned to investigation and compliance needs.
- Measure agent overhead on endpoints, hosts, and cloud workloads.
- Identify duplicate telemetry paths that collect the same event more than once.
- Compare analyst usage against the volume of data being retained.
Security architecture also matters. Centralised logging, distributed tracing, and cloud-native telemetry each create different cost patterns, and there is no universal standard for this yet. Best practice is evolving toward selective, risk-based telemetry design rather than blanket collection. In cyber governance terms, the question is whether observability is still improving decision-making or simply expanding the bill. These controls tend to break down when legacy systems, multi-cloud estates, and unmanaged agents all feed the same pipeline because cost attribution becomes fragmented and optimisation is delayed.
Common Variations and Edge Cases
Tighter telemetry governance often reduces cost, but it can also increase friction for developers, cloud operators, and incident responders, requiring organisations to balance visibility against operational overhead. The trade-off is not always obvious, because a platform may look expensive on paper while still being cheaper than the labour required to stitch together fragmented logs during an incident.
One common edge case is regulated retention. Some teams appear over-instrumented because they are holding data for audit or legal reasons rather than operational preference. Another is bursty environments, where short-lived workloads create sudden ingest spikes that make a healthy platform look unaffordable. In those cases, the issue is often not the observability model itself but the mismatch between workload patterns and collection policy.
There is also a difference between necessary richness and unnecessary duplication. Traces, logs, and metrics serve different purposes, but sending all three at maximum detail everywhere rarely produces proportional value. Where agentic automation is involved, the risk grows further because tool-using systems can create more events, more state changes, and more review requirements. That intersection is still an emerging area, and current guidance suggests treating automated telemetry growth as a capacity and governance problem rather than a pure infrastructure issue. The platform becomes unsustainable when every new service, agent, or control adds cost faster than it adds investigative clarity.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
NIST CSF 2.0 and NIST SP 800-53 Rev 5 set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| NIST CSF 2.0 | DE.CM | Continuous monitoring depends on telemetry that stays affordable enough to sustain. |
| NIST SP 800-53 Rev 5 | AU-2 | Event logging choices determine data volume, cost, and usefulness for investigations. |
Tune monitoring scope and retention so detection value stays high without runaway observability spend.
Related resources from NHI Mgmt Group
- What are the signs that an on premise AI platform is becoming hard to operate safely at scale?
- How can security teams tell whether their CIAM stack is becoming too expensive to govern?
- What breaks when a virtualisation platform becomes too expensive to stay on?
- When does DAST become too expensive to scale effectively?
Deepen Your Knowledge
Reviewed and updated by the NHIMG editorial team on September 1, 2026.
NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org