By NHI Mgmt Group Editorial TeamDomain: Cyber SecuritySource: DataBahnPublished October 8, 2025

TL;DR: Legacy SIEMs struggle with rising telemetry volumes, long retention needs, and the need to keep data both searchable and secure, according to DataBahn. The governance shift is no longer about collecting more logs, but about controlling access, cost, and response speed across the data pipeline.


At a glance

What this is: This is an analysis of why security data lakes are being adopted to centralise logs, improve investigation speed, and reduce SIEM cost pressure.

Why it matters: It matters because identity, access, and retention controls now determine whether security data remains usable for SOC work without becoming a governance or budget bottleneck.

By the numbers:

  • Filtering and enriching telemetry before it reaches the SIEM has reduced data volumes by 50 to 70 percent in production deployments, cutting SIEM licensing costs by more than half without sacrificing the underlying log.
  • One medical device manufacturer running OT-heavy manufacturing sites cut Splunk costs by over 50 percent within seven days of deploying edge-level filtering and enrichment, without dedicating engineering bandwidth to the rollout.

👉 Read DataBahn's analysis of legacy SIEM limits and security data lakes


Context

Security data lakes exist because legacy SIEM architectures were built for a world with less telemetry, fewer environments, and simpler retention demands. As log volume grows across cloud, on-premises, and hybrid estates, the core problem becomes one of access governance as much as storage. If security teams cannot keep raw data searchable, protected, and selectively retrievable, they lose both investigation speed and control over cost.

The identity angle is often overlooked, but it is central. Security data is only useful if the right analysts, hunters, and automation workflows can reach it without broadening edit rights or exposing sensitive telemetry to unnecessary users. In that sense, the article is really about securing a high-value operational dataset, not just storing logs more cheaply.


Key questions

Q: How should security teams secure a security data lake without slowing investigations?

A: Use separate roles for ingestion, search, and administration, and make read access broad enough for investigations but narrow enough to prevent tampering. Add full audit logging for queries, exports, and retention changes. The goal is to keep evidence trustworthy while still making it usable for SOC work and compliance.

Q: Why do legacy SIEMs struggle when telemetry volume keeps rising?

A: Because storage, search, and correlation costs rise faster than the quality of the signal. Once teams ingest everything first and decide later, the SIEM becomes a bottleneck for performance, budget, and analyst time. The result is often delayed detection, selective visibility, and too much effort spent managing pipelines instead of threats.

Q: What breaks when security data is centralised without strong access controls?

A: Analysts may lose trust in the evidence if logs can be edited, deleted, or exported too freely. Centralisation without separation of duties also creates a larger blast radius if an account is compromised. A security data lake only improves governance when identity and entitlement controls are stricter than in the systems it replaces.

Q: What should organisations prioritise first: SIEM tuning or data-lake governance?

A: Governance first, because tuning a costly pipeline does not fix weak access, poor retention design, or untrusted evidence. Once roles, retention tiers, and auditability are defined, SIEM tuning becomes more effective because the pipeline is working with cleaner and more intentional data flows.


Technical breakdown

Why legacy SIEM schemas break at telemetry scale

Legacy SIEMs normalise incoming data into predefined schemas so they can correlate events consistently. That model worked when source diversity and volume were lower, but it becomes brittle when organisations ingest years of logs from cloud, endpoint, network, and SaaS sources. The cost is not just storage. It is schema drift, search latency, and the loss of fidelity when teams compress raw events too early. Security data lakes avoid that by preserving data in more flexible forms for later analysis.

Practical implication: preserve raw telemetry outside the SIEM so schema limits do not force premature data loss.

How enrichment changes the ingestion decision

Enrichment attaches context before routing decisions are made. That can include asset identity, threat intelligence, location, or trust level. Once context exists, the pipeline can distinguish between high-value events that deserve SIEM retention and low-value events that can move to cheaper storage. This is why enrichment and filtering are not separate functions. They are the same control applied earlier in the chain, and earlier control is what changes both cost and detection quality.

Practical implication: put enrichment ahead of retention decisions so routing is based on security value, not raw volume.

Security data lake access control and response speed

A security data lake is only useful if it balances two competing needs: protecting sensitive logs from tampering and making them quickly accessible for threat hunting, forensics, and compliance. That is an access-control problem as much as a data-platform problem. If permissions are too broad, logs become editable or exfiltrable. If permissions are too tight, response time suffers. The architecture therefore needs strong read governance, auditability, and separation between ingestion, analysis, and modification paths.

Practical implication: separate write, read, and admin privileges so analysts can investigate without gaining mutable control over evidence.


Threat narrative

Attacker objective: The objective is to compromise the integrity or availability of security data so detection and investigation lose trustworthiness.

  1. Entry begins when telemetry is concentrated into a central pipeline that stores both sensitive events and analytics data, creating a high-value target for abuse or tampering.
  2. Escalation occurs if the platform does not separate ingestion, read, and edit permissions, allowing attackers or insiders to alter data, obstruct investigations, or expand access.
  3. Impact is achieved when logs become unreliable or inaccessible, weakening detection, forensics, retention compliance, and SOC response speed.

NHI Mgmt Group analysis

Security data governance is becoming an identity problem as much as a storage problem. When telemetry becomes a shared enterprise asset, the question is no longer whether logs exist but who can read, move, or alter them. That makes access segregation, audit trails, and operational privilege boundaries part of the security-data architecture itself. Practitioners should treat the data lake as a controlled evidence system, not a passive archive.

Blast-radius control is the real value proposition behind security data lakes. The architectural win is not just lower SIEM spend. It is the ability to preserve high-value telemetry outside expensive hot-path systems while limiting who can reach the most sensitive records. That supports both SOC velocity and governance, especially where retention rules and incident response requirements overlap. The field is moving toward selective exposure, not universal visibility.

Security data lake adoption exposes a common misconception about centralisation. Centralising data does not automatically improve security unless the access model is more mature than the one it replaces. A shared repository can either reduce operational friction or create a larger single point of failure for trust, depending on how permissions are engineered. The practitioner conclusion is clear: centralisation must be paired with stricter identity and entitlement design.

Agentic AI will intensify the demand for governed security data access. If AI-assisted SOC workflows are to query or summarise telemetry safely, they will need tightly bounded read scopes, provenance controls, and strong separation between retrieval and mutation. That brings identity governance directly into data-platform design. Teams preparing for AI-driven operations should assume machine consumers will multiply access pathways unless policy is enforced at the data layer.

Security data lake programmes should be judged by investigation quality, not storage capacity. The most useful measure is whether analysts can reconstruct incidents faster without loosening control over the underlying evidence. That shifts procurement and architecture discussions away from retention alone and toward controlled usability. Practitioners should evaluate whether their platform improves evidence trust, not just log volume handling.

What this signals

Security data lakes will increasingly be evaluated as governed operational systems rather than passive repositories. That means programme owners should expect scrutiny over who can query telemetry, how long evidence remains trustworthy, and whether machine consumers can widen access without oversight. The architectural boundary between SOC tooling and identity governance is getting thinner.

Controlled evidence fabric: this is the practical direction of travel for security data programmes. The data layer must support fast investigation while preserving provenance, auditability, and least-privilege access for both analysts and automation. Teams that treat the lake as part of the identity plane will be better positioned for AI-assisted detection workflows.

As more security workflows become automated, identity controls around data access will matter more than raw storage scale. Practitioners should prepare for privileged queries, service identities, and AI tools to become formal consumers of telemetry, which means their access must be reviewed, scoped, and logged like any other sensitive entitlement.


For practitioners

  • Separate read, write, and admin privileges for security data Limit who can modify schemas, retention rules, and export paths. Analysts should be able to search and investigate without gaining the ability to change or delete evidence. This is the basic control that keeps a central security dataset trustworthy.
  • Move enrichment ahead of retention decisions Enrich telemetry with asset identity, source reputation, and threat context before deciding whether it belongs in the SIEM or lower-cost storage. That lets you route by security value instead of raw volume and avoid paying full price for low-signal data.
  • Define retention tiers by investigation and compliance need Classify logs into hot, warm, and cold storage based on incident response, forensic, and regulatory use cases. Not every record needs the same query speed, but every tier should remain searchable and governed.
  • Audit who can query sensitive telemetry and how Track human and machine access to security data, including automation jobs, investigation notebooks, and AI-assisted workflows. If a tool can query the lake, it is part of the access model and should be reviewed like any other privileged consumer.

Key takeaways

  • Security data lakes solve a real operational problem, but only when they preserve evidence quality and control access tightly.
  • The strongest economic gains come from enriching telemetry before ingestion, not from storing more data more cheaply.
  • Identity governance now extends into the security data layer, where read, write, and admin separation determines whether the SOC can trust what it sees.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

NIST CSF 2.0, NIST SP 800-53 Rev 5 and CIS Controls v8 set the technical controls, while ISO/IEC 27001:2022 define the regulatory obligations.

FrameworkControl / ReferenceRelevance
NIST CSF 2.0PR.AC-4Access governance is central to protecting centralised security telemetry.
NIST SP 800-53 Rev 5AU-2Audit logging matters when security data itself becomes a controlled asset.
CIS Controls v8CIS-6 , Access Control ManagementAccess control management directly supports governed access to sensitive log data.
ISO/IEC 27001:2022A.8.2Privileged access to log data needs explicit control and review.

Map lake permissions to PR.AC-4 and enforce least privilege for read, write, and admin roles.


Key terms

  • Security Data Lake: A security data lake is a centralised repository for storing large volumes of security telemetry in a queryable form. Unlike a narrow SIEM pipeline, it is designed to keep heterogeneous logs accessible at scale so analysts and automation can correlate identity, endpoint, cloud, network, and application evidence.
  • Telemetry Enrichment: Telemetry enrichment adds context to raw security events before they are searched, routed, or retained. That context can include identity, asset metadata, threat intelligence, or location. In practice, enrichment improves triage decisions, reduces noise, and helps organisations avoid paying premium storage costs for low-value data.
  • Evidence Governance: Evidence governance is the set of controls that keep security records trustworthy, searchable, and protected from unauthorised change. It covers access rights, retention, audit logging, and separation of duties so that logs remain usable for incident response, compliance, and forensic review.
  • Retention tier: A retention tier is lower-cost storage used to preserve data for compliance, audit, and later forensic review. It keeps information available without forcing it into the expensive, analyst-facing part of the pipeline.

What's in the full article

DataBahn's full article covers the operational detail this post intentionally leaves for the source:

  • The architecture behind stream enrichment, including how context is attached before ingestion.
  • The practical reasoning behind routing low-value telemetry to cheaper storage tiers.
  • The vendor's explanation of how security data lakes support AI-ready SOC workflows.
  • The compliance and retention discussion for organisations that must keep logs for multiple years.

👉 The full DataBahn article covers security data lake architecture, enrichment flow, and long-retention use cases.

Deepen your knowledge

NHI Foundation Level course, the industry's only accredited NHI security programme, covers NHI governance, machine identity security, and secrets management. It is designed for practitioners who need to connect identity controls to broader security operations and governance.
NHIMG Editorial Note
Published by the NHIMG editorial team on August 2, 2026.
NHI Mgmt Group — the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org