Security teams should filter at the pipeline and SIEM layers, enrich the remaining data, and tune detections for the environment. The goal is to drop low value or sample findings before they consume storage and licensing, while preserving records that support incident response, threat hunting, and correlation across cloud sources. Cloud native SIEMs make that approach far easier to sustain at scale.
Why This Matters for Security Teams
GuardDuty can surface valuable signals, but a default all-findings feed is rarely economical or operationally useful inside a SIEM. The real issue is not whether the source is trustworthy, but whether every event deserves indexed storage, correlation, and analyst attention. Security teams need to preserve detections that indicate active compromise, privilege misuse, or lateral movement, while avoiding expensive ingestion of repetitive, low-context, or low-actionability findings. That balance aligns well with the NIST Cybersecurity Framework 2.0, which emphasizes outcomes rather than blanket data accumulation.
Practitioners often get this wrong by treating SIEM ingestion as a binary choice between “all logs” and “no logs.” The better approach is selective retention, supported by detection engineering and clear rules for what the SIEM must hold for investigation, threat hunting, and audit. Cost control matters because cloud logging bills often scale faster than the security value they deliver when noisy events are left unfiltered.
In practice, many security teams only discover that their GuardDuty feed is too expensive after licensing pressure or analyst overload has already reduced the quality of detection response.
How It Works in Practice
The most effective pattern is to decide, before ingestion, which GuardDuty findings are operationally essential, which are useful only in raw form, and which can be summarized or dropped. That usually starts with the finding type, severity, affected account, and whether the event is actionable in your environment. For example, detections tied to credential abuse, reconnaissance on sensitive assets, or confirmed malware activity often deserve full retention. Lower-value anomalies may only need aggregated metrics or short-term storage outside the SIEM.
Teams usually implement this in layers:
- Filter at the pipeline so low-value events never reach premium indexing.
- Normalize and enrich events that will be retained with account, asset, and workload context.
- Map retained findings to detections that analysts actually use for triage or correlation.
- Keep a raw archive for forensic access when evidence is needed but not SIEM-indexed.
That design works best when the SIEM is reserved for searchable security-relevant events and not used as a bulk archive. It also helps to define retention by detection purpose. Some findings are needed for real-time alerting, others for weekly hunting, and some only for post-incident reconstruction. Controls for log handling and monitoring in NIST SP 800-53 Rev 5 Security and Privacy Controls are useful here because they separate collection, review, and retention decisions from the raw source stream.
These controls tend to break down in multi-account AWS environments where different business units define “useful” differently and the same finding class has inconsistent operational value across accounts.
Common Variations and Edge Cases
Tighter filtering often reduces SIEM cost, but it also increases the risk of over-pruning, so organisations must balance savings against the chance of losing an early signal. That tradeoff is especially important when GuardDuty is one of several cloud security sources and correlations depend on the exact event chain rather than a single high-severity alert.
Best practice is evolving around what should be retained versus summarized. There is no universal standard for this yet, because acceptable noise levels depend on account maturity, cloud architecture, and the maturity of the detection content already in place. Mature teams often keep more data for crown-jewel workloads and less for sandbox, ephemeral, or low-risk accounts.
Edge cases include environments with strict regulatory retention needs, security operations teams that rely on custom threat hunting queries, and organizations using multiple SIEM back ends with different pricing models. In those cases, the right answer may be to route some findings to a cheaper data lake while preserving only the highest-value records in the SIEM. The key is to treat retention as a detection design decision, not a storage accident.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
NIST CSF 2.0 and NIST AI RMF set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| NIST CSF 2.0 | GV.PO-1 | Policy-driven log retention and filtering supports security outcomes and cost control. |
| NIST AI RMF | AI RMF-style governance is useful for deciding data value and retention boundaries. |
Set logging policy by security value, not raw volume, then enforce it across GuardDuty ingestion paths.
Related resources from NHI Mgmt Group
- How should security teams reduce SIEM ingestion costs without losing detection value?
- How should security teams reduce SIEM cost without losing evidence quality?
- How should security teams reduce SIEM noise without losing important alerts?
- How should security teams reduce shelfware without weakening detection coverage?
Deepen Your Knowledge
Reviewed and updated by the NHIMG editorial team on September 1, 2026.
NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org