By NHI Mgmt Group Editorial TeamDomain: Cyber SecuritySource: DataBahnPublished September 3, 2025

TL;DR: Legacy SIEMs create cost, latency, and context problems because ingestion scales faster than security value, and DataBahn argues that pre-SIEM pipelines, upstream enrichment, and open schemas can reduce noise while making telemetry usable for detection and AI-driven operations. The architectural shift matters because security data now has to be governed before it reaches analytics layers, not after.


At a glance

What this is: This is an independent analysis of why legacy SIEM-centric SOC architecture is struggling, with the key finding that pre-SIEM pipelines and upstream enrichment are now the control point for usable telemetry.

Why it matters: It matters because IAM, NHI, and security teams increasingly depend on telemetry that is contextual, governable, and AI-ready, especially where identity signals, cloud logs, and machine activity must be correlated reliably.

By the numbers:

👉 Read DataBahn's analysis of legacy SIEM limits and security data pipelines


Context

Legacy SIEM programmes often fail because they treat ingestion volume as progress. In practice, more logs can mean less clarity when data is duplicated, delayed, poorly normalised, or missing business context, and that leaves SOC teams paying to store noise instead of improving detection. In security data pipelines, the primary problem is not collection but governance over what telemetry is trusted, enriched, and routed.

The identity angle is real even in a telemetry article: identity, NHI, and workload signals are only useful when they carry enough context to support access review, incident triage, and correlation across systems. That makes pre-SIEM routing relevant to IAM and NHI programmes, not just to SOC engineering. For organisations moving toward AI-assisted operations, AI-ready data becomes a prerequisite rather than an optimisation.

This is a typical enterprise problem, not an edge case. Most large SOCs inherit fragmented logging, multiple schemas, and brittle integrations, then try to compensate with more ingest and more tuning instead of a cleaner data plane.


Key questions

Q: How should security teams reduce SIEM noise without losing important alerts?

A: Focus on context, not volume. Enrich events with identity, location, device, and reputation data before triage so alerts are prioritised by risk rather than by event type alone. This reduces false positives, shortens investigation paths, and helps analysts spend time on evidence instead of manual lookups.

Q: Why do legacy SIEM architectures struggle with modern cloud and identity data?

A: Because they assume a stable log structure and a centralised processing model. Modern cloud, identity, and NHI telemetry arrives in inconsistent formats, at high velocity, and with changing context requirements. A central SIEM then becomes a bottleneck for normalisation, cost, and latency rather than a control point for detection.

Q: What do security teams get wrong about GenAI in the SOC?

A: They often assume the model reduces the need for analyst judgment. In practice, GenAI reduces reading and writing time, but the analyst still owns interpretation, prioritisation, and escalation. If the team uses the model to replace verification, it will amplify mistakes instead of reducing workload.

Q: Should organisations replace SIEMs with security data pipelines?

A: Not necessarily. The better model is to stop treating the SIEM as the centre of the architecture and use the pipeline to decide what should be seen, stored, or suppressed. SIEMs still matter for correlation and investigation, but only after telemetry has been made usable upstream.


Technical breakdown

Why legacy SIEM ingestion becomes a bottleneck

Legacy SIEMs are designed to absorb and index large volumes of telemetry, but that model creates predictable failure points: format rigidity, expensive tuning, delayed detection, and linear cost growth. When every event must be shipped, normalised, stored, and queried centrally, the SOC inherits latency and noise before it can extract signal. The architecture also encourages over-collection because teams assume more data equals better detection, even when the underlying pipeline cannot preserve context or pace with source churn.

Practical implication: move from raw-volume thinking to telemetry governance, with routing and normalisation decisions made before SIEM ingestion.

How upstream enrichment changes security data quality

Upstream enrichment attaches context while telemetry is still in motion, so the pipeline can classify assets, resolve identities, and tag likely risk before events reach analytics tools. This matters because enrichment is what makes filtering intelligent. Without context, a clean internal event and a malicious one may look the same, forcing analysts to spend time reconstructing meaning downstream. Open schemas help by making the data reusable across detection, investigation, and compliance rather than locking it into a single tool format.

Practical implication: enrich at collection and stream layers so the same telemetry can support SOC, IAM, and compliance use cases.

Why AI-ready security data starts in the pipeline

AI tools are only as useful as the telemetry they consume. If logs are inconsistent, noisy, or poorly enriched, reasoning engines inherit ambiguity and produce fragile outputs. Agentic pipelines try to solve that by automating parsing, schema detection, enrichment, and routing in real time. The point is not to add AI to a dashboard, but to make the data fabric structured enough that AI can reason over it reliably. That shifts AI readiness from a product feature to a data architecture requirement.

Practical implication: define AI readiness as a telemetry quality issue and validate the data plane before evaluating AI-driven SOC tools.


NHI Mgmt Group analysis

Legacy telemetry sprawl is now a governance problem, not just an operations problem. Security teams are still treating ingestion, enrichment, and storage as downstream hygiene tasks, but that model breaks once data volumes, vendor churn, and source complexity scale. A SOC cannot trust what it cannot contextualise, and IAM and NHI signals become far less useful when they are delayed or stripped of meaning. The practical conclusion is that telemetry governance now belongs in the security architecture conversation.

Pre-SIEM filtering changes the control point from storage to decision-making. Once enrichment happens before ingestion, routing becomes a security decision, not merely a cost decision. That shift is especially relevant where identity events, service-account activity, and machine-to-machine interactions need to be judged in context rather than preserved indiscriminately. The named concept here is telemetry decisioning: the practice of deciding what telemetry should be trusted, retained, or suppressed before it enters the SIEM. Practitioners should treat that as a core control surface.

AI readiness depends on data quality first and model choice second. Teams keep asking how to use AI in the SOC, but the more important question is whether their data fabric can support reliable reasoning. If schemas are inconsistent and enrichment is deferred, AI merely automates confusion faster. This aligns with NIST CSF and NIST SP 800-53 thinking on protected, monitored, and auditable data flows, and it means the data plane must be governed before AI use cases can be trusted.

Identity signals should be designed for reuse across SOC, IAM, and compliance workflows. The article’s core message extends beyond telemetry engineering because identity and workload events are only useful when they can flow through multiple control layers without rework. That is where open schemas, upstream context, and vendor-agnostic routing matter most. Practitioners should focus on making identity and machine activity legible once, then reusable everywhere the programme needs it.

Vendor-agnostic pipeline architecture is becoming a strategic requirement. As organisations adopt XDR, cloud expansion, and AI-assisted operations, they need a data plane that decouples ingestion from analytics. That reduces lock-in, but more importantly it prevents security operations from inheriting a brittle architecture that cannot evolve with new sources or new detection models. The practical conclusion is to design for change, not for one SIEM’s ingestion contract.

What this signals

Telemetry decisioning: the next SOC control layer is deciding what data deserves full-fidelity retention before it reaches the SIEM. That shift matters because identity and NHI events are only as actionable as the context attached to them, and delaying that context creates both cost and risk. Align the pipeline with identity governance so the data plane can support access review, incident triage, and AI-assisted analysis without rework.

If your organisation is moving toward AI-supported detection or investigation, start by testing whether the telemetry layer can survive schema churn, source volatility, and identity correlation at scale. The practical signal is not how much data you ingest, but whether the data is trusted enough to drive decisions. That is where a resource like Top 10 NHI Issues becomes relevant: telemetry quality and identity control failures usually appear together, not separately.


For practitioners

  • Map telemetry control points before SIEM ingestion Inventory where logs are collected, enriched, filtered, and routed, then document which decisions happen at each stage. Make sure identity, cloud, and NHI events retain enough context to support triage before they hit the SIEM. Use this map to identify where data is being normalised too late.
  • Move enrichment to the edge and stream layers Attach asset, identity, and risk context as close to collection as possible so that downstream tools receive already-classified events. Prioritise upstream enrichment for high-volume sources that currently force analysts to reconstruct meaning after ingestion.
  • Adopt open schemas for reusable telemetry Standardise event structure with a schema that can support detection, investigation, and compliance without repeated transformation. That reduces duplicate normalisation work and makes it easier to correlate identity, workload, and application signals across tools.
  • Test AI readiness against data quality, not marketing claims Before buying AI-driven SOC features, validate whether the underlying telemetry is complete, enriched, and low-noise enough for reasoning engines to use reliably. Poor data quality will undermine any model, no matter how advanced the interface looks.

Key takeaways

  • Legacy SIEM architectures fail when volume is treated as the control objective instead of context, trust, and routing quality.
  • Upstream enrichment and schema normalisation turn telemetry from a storage burden into a decisionable security signal.
  • For identity-heavy environments, the pipeline is now part of governance, because AI and SOC workflows both depend on clean, contextual data.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

MITRE ATT&CK address the attack and risk surface, while NIST CSF 2.0, NIST SP 800-53 Rev 5, CIS Controls v8 and NIST AI RMF set the governance and control requirements practitioners need to meet.

FrameworkControl / ReferenceRelevance
NIST CSF 2.0DE.CM-1Telemetry collection and monitoring are central to the article's SOC architecture critique.
NIST SP 800-53 Rev 5AU-6Log review and analysis depend on enriched, reliable telemetry rather than raw volume.
CIS Controls v8CIS-8 , Audit Log ManagementThe post focuses on how audit data becomes useful or unusable in the SOC.
NIST AI RMFMANAGEAI-ready data and governance are the article's main forward-looking implication.
MITRE ATT&CKTA0007 , Discovery; TA0009 , CollectionThe article is about detecting and analysing telemetry that supports discovery and collection activity.

Map telemetry quality gaps to discovery and collection visibility so detection logic can be tuned upstream.


Key terms

  • Security data pipeline: A security data pipeline is the chain that ingests, filters, enriches, normalises, and routes telemetry before it reaches storage or analytics. In practice, it determines which evidence survives into detection, investigation, and compliance workflows, so it is part of the control environment, not just infrastructure plumbing.
  • Telemetry Decisioning: Telemetry decisioning is the practice of deciding, before SIEM ingestion, what data should be retained at full fidelity, summarised, or suppressed. It combines context, policy, and routing logic so that security teams can manage cost, signal quality, and investigation value at the same time.
  • AI-ready data: Data that is accurate, contextual, and governed enough to support analytics or model decisions without creating avoidable risk. In practice, it means the data carries enough lineage, policy, and ownership information to be trusted, audited, and acted on at the point of use.
  • Pre-SIEM Enrichment: Pre-SIEM enrichment is the process of attaching security context to telemetry before it reaches the SIEM. That context can include identity data, asset ownership, geolocation, or threat intelligence, allowing teams to make a routing decision before they pay indexed-storage costs.

What's in the full article

DataBahn's full blog covers the operational detail this post intentionally leaves for the source:

  • Specific implementation guidance on pre-SIEM routing and enrichment design choices.
  • The mechanics of agentic AI automation across parsing, schema detection, and delivery decisions.
  • Operational detail on reducing ingestion cost without losing investigative fidelity.
  • The company’s framing of how its pipeline model is intended to support future AI use cases.

👉 The full DataBahn post covers the pipeline architecture, enrichment logic, and SOC operating changes in more detail.

Deepen your knowledge

NHI Mgmt Group's NHI Foundation Level course, the industry's only accredited NHI security programme, covers NHI governance, machine identity security, and secrets management. It helps practitioners connect identity controls to the broader security programmes they operate every day.
NHIMG Editorial Note
Published by the NHIMG editorial team on August 2, 2026.
NHI Mgmt Group — the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org