Join our Newsletter — 33% off our NHI Course
Home FAQ Cyber Security What should DevOps and SecOps teams align on…
Cyber Security

What should DevOps and SecOps teams align on before deploying OpenTelemetry at scale?

← Back to all FAQ
By NHI Mgmt Group Editorial Team Updated September 9, 2026 Domain: Cyber Security

They should agree on the metrics that matter, the services to instrument first, and the boundaries for collection and retention. Both teams need a shared dashboard, clear communication channels, and joint decisions on configuration and rollout. That alignment keeps observability useful for operations while avoiding controls that add cost without improving security or reliability.

What DevOps and SecOps need to settle before OpenTelemetry becomes a shared platform

OpenTelemetry is not just an instrumentation choice. At scale, it becomes a shared data pipeline that affects service ownership, incident response, cost, and the quality of security evidence. If DevOps optimises for coverage while SecOps optimises for containment, teams can end up collecting too much low-value data, missing the signals that matter, or creating blind spots through inconsistent deployment rules. The first agreement should be on the business and security questions the telemetry must answer, because that determines where to instrument, what to keep, and what to ignore. The OpenTelemetry project overview at OpenTelemetry is useful background because it frames telemetry as a vendor-neutral standard rather than a single observability product.

Teams also need to decide who owns collector configuration, where trust boundaries sit, and which environments can emit which classes of data. Those choices shape whether telemetry remains operationally useful or becomes an uncontrolled stream of sensitive information. In practice, many teams only discover these disagreements after the first noisy rollout has already inflated storage, delayed troubleshooting, or exposed gaps in auditability.

How a large-scale OpenTelemetry rollout stays useful instead of becoming noisy

The practical model is to treat observability as an architecture decision, not a tooling install. DevOps usually cares about service health, release impact, latency, saturation, and error rates. SecOps usually cares about trace integrity, detection value, data minimisation, retention, and whether telemetry can support investigation without introducing unnecessary exposure. Those goals overlap, but they are not identical, so the rollout plan has to define which signals are mandatory, which are optional, and which are prohibited in shared pipelines.

Good alignment starts with a small set of instrumented services that are representative of the production estate. That allows both teams to validate whether the collected spans, logs, and metrics are actually actionable before scaling across every service. It also forces a decision on the practical boundaries of collection. For example, some attributes may be essential for troubleshooting but inappropriate for long-term storage, while others may be useful for security correlation but too expensive to retain at high volume.

  • Define the minimum telemetry schema for operational and security use cases before expanding coverage.
  • Agree on collector ownership, change control, and escalation paths for broken or unsafe configurations.
  • Set retention and access rules based on the data class, not on a one-size-fits-all default.
  • Test whether the shared dashboard helps both incident triage and service health review, rather than serving only one team.

That working model also needs a feedback loop. If the first deployment shows that certain traces are too sparse to diagnose failures or too rich to store economically, the schema or sampling approach should be adjusted before wider rollout. OpenTelemetry becomes most valuable when teams treat its output as governed operational evidence, not as an infinite telemetry feed. The guidance breaks down when each team optimises its own pipeline independently and no one is accountable for the full data lifecycle.

Where OpenTelemetry rollout decisions become expensive or risky

Tighter telemetry control often improves security and cost discipline, but it can also reduce visibility if teams over-correct and sample away the wrong events. That tradeoff matters most when the same data must support both performance troubleshooting and security investigation, because the collection strategy that is cheap for one use case may be weak for the other. Guidance is still evolving on the best balance between broad coverage and data minimisation, so organisations should treat the retention model as a governed decision rather than an assumption.

Another edge case is multi-team ownership. If platform teams own the collector but application teams own the instrumentation, gaps can appear at service boundaries unless the interface is explicitly defined. The other common failure is treating telemetry labels and attributes as harmless metadata when they may actually reveal tenant details, user context, or environment structure that should not be broadly accessible. For that reason, the collection boundary should be reviewed as carefully as the dashboard design.

Teams should also distinguish between what they need for live operations and what they need for forensic review. Those requirements are related, but they often justify different retention periods, access controls, and sampling choices. When a rollout does not make that distinction explicit, the system tends to drift toward either overcollection or underuse.

Risk and Threat Considerations

At scale, OpenTelemetry creates a material exposure if teams treat observability data as low-risk operational noise. Telemetry can contain service maps, request metadata, deployment timing, authentication context, and other clues that improve troubleshooting but also increase sensitivity, retention burden, and access risk. The security concern is not the protocol itself, but the way uncontrolled instrumentation can expand the amount of information available to too many people or systems.

Failure mechanism: Risk materialises when collection scope, sampling, and retention are decided locally by separate teams, then propagated into shared pipelines without a common policy. That produces inconsistent data quality, accidental capture of sensitive attributes, and weak accountability over who can read or export telemetry. In an incident, the same sprawl can slow investigation because the useful signals are diluted by noise while the truly relevant spans were never standardised.

Impact: The organisation can lose diagnostic value, increase storage and processing costs, and create avoidable confidentiality exposure in logs, traces, or metrics. It can also undermine incident response if investigators cannot trust whether critical events were sampled, retained, or masked consistently across services.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

MITRE ATT&CK address the attack surface, CIS Controls v8 and NIST CSF 2.0 set the technical controls, and ISO/IEC 42001:2023 define the regulatory obligations.

FrameworkControl / ReferenceRelevance
CIS Controls v88 — Audit Log ManagementOpenTelemetry is a logging, tracing, and monitoring pipeline.
Recommendation — Define and protect telemetry collection, retention, and review rules for production environments.
NIST CSF 2.0DE.CM-1 — Monitoring for anomalies and eventsShared telemetry underpins detection and operational monitoring.
PR.PT-1 — Audit/log recordsTelemetry rollout requires agreed logging and trace-record handling.
Recommendation — Align telemetry coverage to the events your monitoring process must actually detect. Establish consistent logging and trace-record practices before scaling instrumentation.
MITRE ATT&CKT1112 — Modify RegistryTelemetry and collector controls can be tampered with if change paths are weak.
Recommendation — Hunt for unauthorized collector or instrumentation changes that alter observability data.
ISO/IEC 42001:2023A.6 — AI system lifecycleUse only if telemetry is feeding AI/agent workflows; otherwise not central enough.
Recommendation — Govern telemetry inputs as controlled lifecycle data when they feed AI-enabled operations.

Practitioner Guidance

What to prioritise: lock down the first-use cases before broad instrumentation. If DevOps and SecOps cannot name the decisions the telemetry must support, they should not expand collection beyond a pilot set of services.

What to verify: confirm that collector ownership, data boundaries, and retention rules are defined for each environment, not implied by platform defaults. The control is only real when both teams can show the same answer to “what data is collected, where does it go, and who can use it?”

What good looks like: one shared dashboard supports service reliability and security triage, the telemetry schema is stable enough for trend analysis, and changes to instrumentation follow the same approval path as other production changes.

Practitioner takeaway: the biggest mistake is scaling collection before scaling agreement; once telemetry becomes infrastructure, misalignment turns into cost, noise, and blind spots that are harder to unwind than to prevent.

Deepen Your Knowledge

Sign up to our weekly newsletter — get 33% off our NHI Foundation Level Course

    NHIMG Editorial Note
    Reviewed and updated by the NHIMG editorial team on September 9, 2026.
    NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org