Subscribe to the Non-Human & AI Identity Journal
Home FAQ Cyber Security What do organisations get wrong about postmortem and…
Cyber Security

What do organisations get wrong about postmortem and root cause analysis?

← Back to all FAQ
By NHI Mgmt Group Editorial Team Updated August 2, 2026 Domain: Cyber Security

They treat RCA as an after-action report instead of a control loop. A useful postmortem should change playbooks, ownership, training, and escalation criteria. If the review does not alter future response behaviour, it only records the failure without reducing the next one.

Why This Matters for Security Teams

Postmortem and root cause analysis are meant to reduce repeat failure, yet many organisations still treat them as documentation exercises. That misses the operational purpose of learning reviews: identifying where detection, response, escalation, or ownership failed and then changing the control environment. NIST’s NIST SP 800-53 Rev 5 Security and Privacy Controls is useful here because it frames response and improvement as managed security functions, not optional follow-up tasks.

The common mistake is stopping at a narrative of events. Teams capture timelines, interview responders, and document technical cause, but they do not convert findings into revised alerting thresholds, access boundaries, approval rules, or recovery steps. That gap matters because the same weak point often remains in production after the report is closed. In environments with shared services, cloud automation, or identity-heavy workflows, the issue may sit in a dependency chain rather than the obvious incident trigger.

Practitioners also underestimate how often “root cause” is multiple causes working together, such as poor telemetry, unclear escalation ownership, and delayed decision-making. In practice, many security teams encounter the true failure only after the same issue has repeated, rather than through intentional learning and control change.

How It Works in Practice

A useful RCA process starts with a bounded question: what failed, what should have detected it, who was expected to act, and what control gap allowed the outcome? That framing keeps the review focused on operational improvement instead of blame. The strongest postmortems usually separate symptom, contributing factors, and control failures, because those are not the same thing. A failed service, a missed alert, and a delayed approval may all appear in the same incident, but each requires a different correction.

For security teams, the output should map directly to control owners and change mechanisms. That usually means updating incident runbooks, refining SIEM or SOAR logic, revising access policies, and retraining responders. If identity or privileged access played a role, the review should also examine whether PAM, JIT access, or privileged session logging was missing, misconfigured, or bypassed. Where the issue involves secrets, API keys, or service accounts, the corrective action needs to address rotation, scoping, and storage, not just awareness.

Current guidance suggests the most effective reviews tie each finding to a measurable change. A practical pattern is:

  • identify the failure mode and the control that should have stopped it;
  • assign a single accountable owner for remediation;
  • set a deadline and verify the fix in production or a realistic test environment;
  • confirm the change has altered alerts, approvals, or recovery steps;
  • record what remains unresolved and why.

That approach aligns well with CISA’s Known Exploited Vulnerabilities Catalog style of prioritisation, which emphasizes action on real exposure rather than theoretical completeness. These controls tend to break down when incident data is fragmented across cloud, endpoint, and identity teams because no single owner can see the full chain of failure.

Common Variations and Edge Cases

Tighter postmortem discipline often increases coordination overhead, requiring organisations to balance speed of recovery against depth of analysis. That tradeoff is real, especially during high-severity incidents or customer-facing outages. The best practice is evolving, but the baseline remains consistent: the review must change something in the environment, or it is not a control improvement.

One edge case is recurring low-severity incidents. Teams sometimes dismiss them because each event looks minor, yet the repetition often signals a systemic issue in threshold tuning, ownership, or backlog handling. Another is incidents with incomplete telemetry. In those cases, current guidance suggests stating uncertainty explicitly rather than inventing certainty; a partial RCA is still useful if it identifies missing logs, broken traces, or absent audit data as findings in themselves.

Identity-related failures deserve special care. If a compromised credential, mis-scoped service account, or over-privileged role contributed to the incident, the postmortem should check whether access governance, privilege review, and credential lifecycle controls were actually enforced. For organisations operating automated pipelines or agentic systems, the same logic applies to machine identities and tool permissions. The review should ask whether the agent had too much execution authority, whether its actions were observable, and whether a human approval step was appropriate for the risk.

Where the environment mixes legacy tooling, outsourced operations, and cloud-native controls, the answer can be messy rather than elegant. In those settings, the real failure is often not the incident itself but the absence of a reliable mechanism to turn lessons into enforced change.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

OWASP Non-Human Identity Top 10, OWASP Agentic AI Top 10 and MITRE ATT&CK address the attack and risk surface, while NIST CSF 2.0 and NIST AI RMF set the governance and control requirements practitioners need to meet.

FrameworkControl / ReferenceRelevance
NIST CSF 2.0RS.IMPost-incident improvements must become repeatable response improvements.
NIST AI RMFGOVERNAI governance principles apply when reviews cover agentic or AI-driven operations.
OWASP Non-Human Identity Top 10RCA often exposes machine identity and secret lifecycle failures.
OWASP Agentic AI Top 10Agent tool permissions and unsafe actions can be root causes in automated environments.
MITRE ATT&CKT1078Valid account abuse is a frequent incident contributor that RCA should expose.

Verify agent permissions, logging, and human approval steps after any autonomous action failure.

NHIMG Editorial Note
Reviewed and updated by the NHIMG editorial team on August 2, 2026.
NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org