TL;DR: Security teams are being pushed to ingest more telemetry for AI-driven operations, but noisy, inconsistent data still creates analyst overhead instead of better decisions, according to Anomali. The real productivity gain comes from engineering signal through normalization, context enrichment, and workflow-driven analytics, not from feeding models everything.
NHIMG editorial — based on content published by Anomali: Data Hygiene for AI Security: Stop Ingesting Everything, Start Engineering Signal
By the numbers:
- Microsoft reports that it processes 78 trillion security signals per day as part of its threat landscape insights.
- In the SANS 2024 SOC Survey, "Too many alerts that we can’t look into/lack of correlation between alerts" appears as an explicit barrier to full SOC capability utilization.
Questions worth separating out
Q: How can AI help SOC analysts without creating more noise?
A: AI helps when it operates on enriched, governed intelligence rather than raw telemetry.
Q: Why do unnormalized logs undermine AI-driven security analytics?
A: Unnormalized logs use different field names, event structures, and semantics across systems, so AI cannot reliably compare them or explain what they mean.
Q: What do security teams get wrong about data volume and visibility?
A: Teams often assume that more telemetry automatically means better visibility, but data volume without purpose just increases noise.
Practitioner guidance
- Classify telemetry by operational purpose Separate compliance retention feeds from detection and response feeds before expanding ingestion.
- Normalize identity and asset fields first Create a standard schema for user, role, privilege, host, application, and environment fields so logs from IAM, endpoint, cloud, and SIEM sources can be compared reliably.
- Attach context that reduces analyst uncertainty Enrich high-value events with role, peer group, access pattern, and asset criticality so alerts reflect decision points instead of raw anomalies.
What's in the full article
Anomali's full article covers the operational detail this post intentionally leaves for the source:
- Examples of how telemetry normalization changes the quality of AI-assisted SOC workflows
- Additional practitioner discussion on data segmentation for detection, retention, and enrichment
- More detail on how identity and asset context improve correlation and triage decisions
- The webinar conversation and leadership commentary that informed the article's operating model
👉 Read Anomali's analysis of data hygiene for AI security and SOC productivity →
AI-ready SOC data hygiene: what security teams are missing?
Explore further
Data hygiene is now an AI governance control, not a housekeeping task. Security teams that feed unstandardized telemetry into AI are effectively delegating judgment to a system that cannot reliably interpret the evidence. That increases the chance of false confidence, wasted analyst effort, and misprioritized response. The control question is no longer whether AI can process more data, but whether the organisation has defined the signal well enough for AI to be trusted with it. The practitioner conclusion is simple: governance starts with data quality.
A question worth separating out:
Q: How can organisations tell whether AI SOC ROI is actually improving?
A: Watch for sustained gains in MTTR, MTTD, alert coverage, and false positive reduction, not just a one-time spike after rollout. Pair those metrics with auditability of the investigation output and with analyst feedback on decision quality. If the numbers improve but trust falls, the model is not healthy.
👉 Read our full editorial: Data hygiene is the missing control in AI-enabled SOC operations