Join our Newsletter — 33% off our NHI Course
Home FAQ Cyber Security What are the signs that an OpenTelemetry setup…
Cyber Security

What are the signs that an OpenTelemetry setup is not giving teams useful observability?

← Back to all FAQ
By NHI Mgmt Group Editorial Team Updated September 8, 2026 Domain: Cyber Security

Common warning signs include noisy dashboards, slow alerting, missing trace coverage for critical requests, and dashboards that show data but do not help isolate bottlenecks. If logs, metrics, and traces are not correlated, teams often spend too long reconstructing incidents. Another signal is when instrumentation growth starts to affect service latency or resource consumption.

When OpenTelemetry Becomes Visibility Without Diagnosis

OpenTelemetry is useful only when it shortens the path from symptom to cause. If teams can collect telemetry but still cannot answer basic questions about request flow, dependency failure, or latency spikes, the setup is not delivering observability in practice. The problem is usually not the absence of data, but the absence of usable context, consistent coverage, or a shared way to interpret what the data means. The NIST SP 800-53 Rev 5 Security and Privacy Controls baseline is relevant here because observability quality depends on control over logging, monitoring, and configuration discipline, not just on collecting more telemetry.

In practice, many teams discover this only after an incident has already forced them to manually reconstruct the request path from fragmented signals.

How Teams Can Tell the Telemetry Is Not Pulling Its Weight

Useful observability shows up as fast, confident investigation. Unhelpful observability shows up as dashboards that are busy but indecisive. When engineers still need to jump across tools, manually match timestamps, or infer service relationships from memory, the telemetry model is not supporting diagnosis. That often means the data exists, but the instrumentation strategy is inconsistent across services, environments, or request types.

A strong test is whether a common production question can be answered without guesswork: which request path failed, which dependency slowed down, whether the issue is isolated or systemic, and whether the symptom is new or recurring. If traces do not cover the important transactions, if metrics are detached from logs, or if correlation identifiers are missing or unreliable, the platform may be producing signals rather than observability.

  • Dashboards report activity but do not reveal where time is spent or why errors cluster.
  • Traces exist for some paths, but not for the business-critical flows that matter most.
  • Logs contain details, but they are not consistently linked to trace or metric context.
  • Alerting fires on volume or thresholds, yet does not help isolate root cause quickly.
  • Instrumentation introduces noticeable overhead, which can distort the very latency being measured.

OpenTelemetry works best when instrumentation is deliberately scoped to the services and transactions that matter most, and when the team can explain what each signal is meant to answer. If the setup grows faster than the organisation’s ability to keep naming, sampling, and context consistent, the data lake can become harder to use than the original problem. This guidance breaks down when teams treat telemetry collection as the same thing as operational understanding.

Where OpenTelemetry Deployments Drift Away From Real Observability

Tighter instrumentation often increases operational overhead, so teams have to balance depth of coverage against noise, cost, and runtime impact.

One common edge case is partial success: the setup works well for a subset of services, but not for the transactions that fail most often or carry the highest business impact. That creates a false sense of maturity because the platform looks healthy on paper while critical gaps remain in the busiest or most fragile paths. Another is over-instrumentation, where teams add too many low-value signals and make it harder to find the meaningful ones.

Guidance versus consensus matters here. There is broad agreement that correlation across logs, metrics, and traces improves diagnosis, but there is no universal consensus on the exact telemetry mix every system should expose. The right balance depends on architecture, request volume, regulatory pressure, and how quickly teams need to diagnose incidents. If the organisation cannot explain why a metric exists, what decision it supports, or how it reduces time to resolution, it is probably accumulating telemetry rather than observability.

Another important edge case is platform sprawl. Multiple teams may adopt different conventions for span names, tags, sampling, or log structure, which makes the aggregated view less reliable over time. In those environments, the observability issue is often governance, not tooling. The setup stops being useful when each team can emit data, but no one can depend on it being comparable.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

MITRE ATT&CK address the attack and risk surface, while NIST CSF 2.0 and CIS Controls v8 set the governance and control requirements practitioners need to meet.

FrameworkControl / ReferenceRelevance
NIST CSF 2.0DE.CM-1 — Monitoring for unauthorized access and anomalous activityObservability quality depends on effective monitoring signals and anomaly visibility.
DE.CM-8 — Vulnerability scans are performedInstrumentation overhead and coverage gaps can mask or distort operational weakness detection.
Recommendation — Validate monitoring coverage so teams can spot and triage abnormal service behaviour quickly. Review telemetry for blind spots that hide service degradation or failure conditions.
CIS Controls v88.2 — Audit Log ManagementLogs, traces, and metrics must be usable and correlated to support diagnosis.
8.6 — Centralized Audit Log ManagementCentralising telemetry only helps if it remains searchable, linked, and operationally useful.
Recommendation — Standardise audit and telemetry records so investigators can correlate events across systems. Centralise telemetry in a form that preserves context for incident analysis.
MITRE ATT&CKT1005 — Data from Local SystemUseful observability can expose local diagnostic data that helps reveal process and failure state.
T1036 — MasqueradingMisleading or low-signal telemetry can obscure what is actually happening in production.
Recommendation — Hunt for whether critical telemetry is missing from the data sources you rely on. Check whether telemetry naming and tagging are obscuring the true service or request path.

Practitioner Guidance

What to prioritise: Focus first on the requests and dependencies that are most expensive to troubleshoot, not on making every service equally observable. The most valuable telemetry is the kind that reduces incident triage time for the paths that matter.

What to verify: Validate that a real production incident can be traced from symptom to dependency to code path without manual stitching across tools. If that requires tribal knowledge, the observability model is not yet serving the team.

Common mistake: Treating more dashboards, more spans, or more alerts as proof of maturity. Volume does not equal usefulness if the signals do not answer a specific operational question or if the overhead starts to degrade service behaviour.

What good looks like: Engineers can quickly identify whether an issue is in application logic, downstream dependency behaviour, or instrumentation blind spots, and they can do so using consistent evidence rather than post-incident reconstruction.

Practitioner takeaway: OpenTelemetry is only useful when it makes decisions faster and investigation narrower; if it mainly increases data production, the organisation has built telemetry collection, not observability.

Deepen Your Knowledge

Sign up to our weekly newsletter — get 33% off our NHI Foundation Level Course

    NHIMG Editorial Note
    Reviewed and updated by the NHIMG editorial team on September 8, 2026.
    NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org