Look for recovery steps that only one or two people can explain, incident decisions that are not written down, and repeated pauses while teams wait for a known expert. Those are signs that the organisation has hero knowledge instead of durable process, which means resilience is not yet scalable.
What the warning signs look like in practice
Resilience becomes too dependent on individual experts when recovery knowledge lives in people rather than in repeatable operating practice. The strongest signals are not abstract: they show up as undocumented workarounds, verbal-only decision paths, and a steady pattern of “ask X” whenever something unusual happens. That creates a fragile organisation because continuity depends on availability, not capability.
A useful way to read the warning signs is to separate healthy expertise from dependency. Healthy expertise shortens investigation and improves judgement; dependency appears when the same person is needed to start, unblock, or finish routine recovery tasks. If a team cannot explain why a step works, who owns it, or what happens when the expert is absent, the resilience model is already concentrated too narrowly.
Look for repeated recovery pauses, especially during incidents, change windows, or failover events. Those pauses usually mean the process is not executable from documentation, tooling, or training alone, and that the organisation is relying on memory under pressure. NIST Cybersecurity Framework 2.0 is useful here because it frames resilience as an operating capability across govern, identify, protect, detect, respond, and recover, not as a single expert trait.
Where individual-expert dependency becomes a resilience failure
The operational failure mode is concentration risk. When one or two people hold the only reliable map of dependencies, rollback order, exception handling, or manual recovery steps, the organisation loses recoverability the moment those people are unavailable. This is especially visible after role changes, vacations, or handovers, when teams discover that “known good” procedures were never written down in a way others can use.
Another common sign is hidden exception handling. Teams may have a procedure for the standard case, but real recovery depends on private tribal knowledge about which alerts to trust, which systems can be restarted safely, or which dependencies must be bypassed first. That gap is not just an efficiency issue, it means the resilience posture cannot scale beyond the original experts. NIST AI Risk Management Framework is a useful reference for the broader governance lesson: resilience depends on documentation, roles, and accountable processes that reduce over-reliance on any one operator, whether the system is AI-related or not.
Teams also tend to overestimate resilience when senior experts are highly responsive. Fast expert response can mask weak process design for a long time. The test is whether another competent operator can make the same recovery decision from the artifacts available at the time, not whether the preferred expert can answer quickly on chat.
How to tell whether the organisation has hero knowledge instead of durable process
The clearest indicator is repeatability. If the same incident class produces the same questions every time, but the answers still depend on who is on call, then the process has not matured into a stable control. Durable resilience has written decision points, testable runbooks, and handover evidence that survive personnel changes.
Practitioners should also watch for mismatch between ownership and actual execution. If one expert is informally doing incident command, change approval, recovery validation, and post-incident review, the organisation has concentrated too many critical functions in one person. That may feel efficient in the short term, but it creates a single point of failure in the human layer. NIST SP 800-53 Rev 5 supports this interpretation through controls on access, auditability, configuration management, and contingency planning, all of which become weaker when knowledge is not distributed.
The practical question is whether the team can demonstrate recovery without live expert coaching. If the answer is no, the organisation should treat that as a resilience defect, not a training inconvenience. At that point the issue is not simply “one person is very good”, it is that the operating model cannot be trusted to hold under absence, turnover, or simultaneous incidents.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
NIST CSF 2.0 and NIST SP 800-53 Rev 5 set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| NIST CSF 2.0 | RC.RP — Recovery Planning | Recovery depends on repeatable processes, not one expert. |
| GV.OC — Organizational Context | Defines ownership and accountability for resilience operations. | |
| Recommendation — Test whether recovery steps work without the original expert present. Assign explicit ownership for recovery knowledge and execution. | ||
| NIST SP 800-53 Rev 5 | CP-2 — Contingency Plan | Contingency plans reduce reliance on undocumented expert memory. |
| AU-12 — Audit Record Generation | Recorded decisions prevent recovery knowledge from living only in people. | |
| CM-2 — Baseline Configuration | Baselines support recoverability without ad hoc expert workarounds. | |
| Recommendation — Document and exercise contingency steps for critical services. Capture key recovery decisions in auditable records. Maintain baselines so operators can restore known-good states. | ||
Practitioner Guidance
What to verify: Confirm whether recovery steps are actually executable from current runbooks, tickets, diagrams, and automation, or whether the team is depending on remembered shortcuts. The most telling test is a supervised recovery drill with the named expert absent.
Common mistake: Treating a fast expert escalation path as resilience. Speed of access to expertise is not the same as distributed capability, and it often hides fragile process design until an outage, holiday, or resignation exposes it.
What good looks like: A competent on-call operator can perform the critical recovery sequence from written material, understand where the risky judgement calls are, and complete handover without needing the originator of the process in the room.
Practitioner takeaway: Resilience is too dependent on individual experts when the organisation cannot recover safely without them, even if those experts are technically excellent. The goal is not to remove expertise, but to convert it into shared, testable, and survivable operating knowledge.
Related resources from NHI Mgmt Group
Deepen Your Knowledge
Free weekly newsletter
Subscribe to the NHI & AI Identity Journal
The latest on NHI and Agentic AI security – articles, research, breaches, news and events every week.
Bonus 33% off our NHI Course when you subscribe.
Reviewed and updated by the NHIMG editorial team on October 11, 2026.
NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org