Measure and improve is the feedback loop used after a failure test. Teams observe how the platform behaved, identify where it succeeded or failed, then adjust configuration, automation, documentation, or architecture to strengthen recovery and reduce future incident impact.
What Measure and Improve Actually Means
Measure and improve is the post-test feedback loop that turns a failure test into learning. Teams observe how the platform behaved, compare actual recovery to the intended outcome, and then refine controls, automation, documentation, or architecture so the next incident is less damaging.
This step matters because failure tests are only useful when they change the system. Without measurement, organisations cannot tell whether a recovery path is fast, repeatable, and complete, or merely successful once under ideal conditions.
Where It Fits in Resilience Testing
The phrase usually appears after a chaos exercise, game day, tabletop, or controlled failure injection. The purpose is not to prove that systems never fail, but to expose how recovery really works under pressure, including what operators can see, what the platform can self-heal, and where manual intervention is still required.
That makes the “measure” part as important as the “improve” part. Good measurement captures recovery time, failed assumptions, dependency delays, alert quality, and whether the runbook matched reality. The improvement stage then converts those observations into changes that actually reduce future incident impact rather than merely documenting lessons learned.
What Gets Measured and Changed
Useful measures include time to detect failure, time to restore service, the number of steps needed for recovery, whether failover was automatic, and whether the team had the access and context needed to act quickly. NHI Mgmt Group’s Ultimate Guide to Non-Human Identities notes that only 5.7% of organisations have full visibility into their service accounts, a reminder that weak visibility can make recovery slower and less reliable when automation depends on machine access.
The changes that follow are usually practical: tuning alerts, fixing broken runbooks, removing brittle dependencies, tightening permissions that block recovery, or redesigning components that fail in predictable ways. In mature resilience programmes, the output is not just a report, but a concrete reduction in blast radius and restoration friction.
Why This Loop Changes the Security Posture
Measure and improve strengthens resilience because it closes the gap between theoretical controls and observed behaviour. A system can look well defended on paper yet still recover badly if alerts are noisy, dependencies are opaque, or the recovery process relies on tribal knowledge.
It also exposes control weaknesses that only appear during failure. A recovery task may depend on credentials, automation, or third-party services that work in normal operations but slow the team down during an outage. Measuring those failure modes helps teams decide whether the control needs redesign, better instrumentation, or different ownership.
Risk and Threat Considerations
Resilience gaps become security risks when a team cannot see how a system behaves during failure or cannot restore it fast enough. The main danger is not only longer downtime, but also hidden control failures that leave recovery paths brittle, inconsistent, or dependent on manual heroics.
Failure mechanism: Poor measurement leaves organisations with incomplete evidence about restoration time, dependency failure, and operator friction, so weak recovery paths stay in place until a real incident exposes them.
Impact: Recovery becomes slower and less predictable, which increases outage duration, operational disruption, and the chance that a security incident or service failure cascades into a larger business event.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
NIST CSF 2.0 and NIST SP 800-53 Rev 5 set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| NIST CSF 2.0 | RC.RP-01 — Recovery Plan Execution | Measure and improve operationalises recovery by comparing actual restoration to the intended recovery plan. |
| RC.IM-01 — Improvements | This term is the improvement loop that turns failure-test results into control and process changes. | |
| GV.RR-02 — Roles, Responsibilities, and Authorities | Improvement after a failure test depends on clear ownership for remediation and recovery decisions. | |
| Recommendation — Use RC.RP-01 to test whether recovery actions restore services as designed and refine weak steps. Use RC.IM-01 to capture test findings and update resilience controls, runbooks, and architecture. Assign owners under GV.RR-02 so test findings are converted into tracked remediation work. | ||
| NIST SP 800-53 Rev 5 | CP-4 — Contingency Plan Testing | This control requires testing contingency arrangements and using results to strengthen recovery capability. |
| Recommendation — Use CP-4 to exercise recovery plans and update them based on observed test results. | ||
Practitioner Guidance
Why practitioners should care: Treat the output of a failure test as an engineering input, not a postmortem artifact. The real value of this term is that it turns observation into durable changes in configuration, automation, and architecture.
Common misunderstanding: A successful test does not mean the system is resilient if the team needed extra context, manual workarounds, or privileged intervention to recover. The test only proves what happened under those specific conditions.
Practitioner takeaway: If a failure test does not produce at least one measurable improvement, the loop is incomplete.
Related resources from NHI Mgmt Group
- How should teams measure whether coding agent rules actually improve task accuracy and tool use?
- How should security teams measure the business value of identity security?
- How can SOC teams use identity context to improve response to agent activity?
- How should organisations measure identity security ROI beyond license savings?
Deepen Your Knowledge
Reviewed and updated by the NHIMG editorial team on September 24, 2026.
NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org