Join our Newsletter — 33% off our NHI Course
Home› FAQ› Governance, Ownership & Risk› What should teams do immediately after finding gaps…
Governance, Ownership & Risk

What should teams do immediately after finding gaps in infrastructure as code coverage?

← Back to all FAQ
By NHI Mgmt Group Editorial Team Updated October 10, 2026 Domain: Governance, Ownership & Risk

Bring manual or hidden resources into source control, then reconcile the declared state with production before the next outage. Any component that only exists in a console or in tribal knowledge is a recovery liability because it cannot be reproduced consistently under pressure.

Why IaC Coverage Gaps Become Recovery Problems

Infrastructure as code is not just a deployment preference, it is the record that lets teams rebuild, verify, and compare what should exist against what does exist. When coverage is incomplete, the gap is often hidden until an incident, rollback, or scaling event forces a rebuild. At that point, the missing item becomes a recovery dependency instead of a convenience issue.

The practical problem is drift between declared and actual state. If a component only exists in a console, in a one-off script, or in someone’s memory, it can be missed during replacement, recreated inconsistently, or left behind during teardown. That makes the environment harder to reason about and increases the chance that the next change, outage, or failover exposes a configuration that was never under normal control.

Teams should treat any uncovered resource as evidence that the system is not yet fully reproducible. The immediate goal is not perfect aesthetic alignment in the repository, but closing the gap on the resources that matter most to service restoration, access, dependency order, and configuration consistency.

What “bring it into source control” should mean in practice

Bringing a manual or hidden resource into source control means capturing the real operational state in a versioned, reviewable form so the team can recreate it without tribal knowledge. For infrastructure work, that usually includes the resource definition, its dependencies, its variables or parameters, and the ownership needed to maintain it safely over time.

This step only works if the code reflects production truth, not an idealised version of it. Teams should inspect the live environment, compare it with the intended model, and reconcile differences before assuming the repository is authoritative. If the code is added without that reconciliation, the team may simply preserve the mismatch in a cleaner format.

A good rule is that anything required to restore service, satisfy security controls, or avoid an outage should be treated as part of the managed baseline. Anything outside that baseline should be either intentionally exempted with a documented reason or moved into the same control plane as the rest of the estate.

Reconcile declared state with production before the next outage

Once the gap is identified, the next step is to align the declared state and the live environment while the system is still healthy. That often means checking resource settings, permissions, network paths, secrets references, and dependency order, then deciding whether the source of truth should be updated to match reality or the environment should be corrected to match the intended configuration.

For teams operating at scale, the important distinction is between a one-time fix and an ongoing control. A one-time export into source control helps, but durable improvement comes from preventing the same class of drift from reappearing through reviews, automated checks, and change discipline. A repository that still allows unmanaged console edits will drift again.

Reconciliation should also account for recovery behaviour. If a component is needed during failover, rebuild, or incident containment, teams should verify that it can be recreated from code, that its dependencies are also codified, and that there is no hidden manual step that only works when the original operator is available.

Risk and Threat Considerations

Coverage gaps create operational exposure because undocumented infrastructure is harder to recreate, audit, and secure. They also widen the blast radius of a mistake, since a single missing definition can block recovery or leave an unmanaged path in place long after the rest of the estate has been controlled.

Failure mechanism: hidden or manually created resources accumulate outside the change system, so restoration depends on memory, console access, or ad hoc fixes instead of repeatable infrastructure.

Impact: recovery becomes slower and less reliable, drift becomes harder to detect, and an outage can become a prolonged service interruption if the missing resource was part of the critical path.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

NIST SP 800-53 Rev 5 and CIS Controls v8 set the governance and control requirements practitioners need to meet.

FrameworkControl / ReferenceRelevance
NIST SP 800-53 Rev 5CM-2 — Baseline ConfigurationIaC coverage gaps are baseline gaps that must be managed.
CM-6 — Configuration SettingsReconciliation depends on comparing live settings with the intended configuration.
CM-8 — System Component InventoryHidden resources are an inventory and accountability problem.
Recommendation — Inventory unmanaged infrastructure and turn it into a controlled baseline. Standardize and verify configuration settings against the declared state. Maintain a complete inventory of infrastructure components and owners.
CIS Controls v8CIS-4 — Secure Configuration of Enterprise Assets and SoftwareIaC coverage gaps expose unmanaged configuration drift and hidden assets.
CIS-5 — Account ManagementManual resources often persist through unmanaged access paths and ownership gaps.
Recommendation — Harden and codify configurations so live systems match approved builds. Reconcile access ownership for any resource brought under source control.

Practitioner Guidance

What to prioritise: start with the resources that affect recovery first, especially networking, access dependencies, stateful services, and any control-plane component that would delay rebuild or failover if it disappeared. Less critical components can follow once the restoration path is reproducible.

What to verify: confirm that the repository can describe the resource well enough for a second team member to recreate it without asking the original operator. If that is not true, the gap is still operationally material even if the system appears stable today.

Common mistake: teams often stop after importing the obvious resource and ignore its attached assumptions, such as permissions, secrets references, or environment-specific overrides. That leaves the real recovery problem untouched.

Practitioner takeaway: the immediate objective is to eliminate unreproducible state, because anything that cannot be rebuilt from code is already an incident risk, even before it fails.

Free weekly newsletter

Subscribe to the NHI & AI Identity Journal

The latest on NHI and Agentic AI security – articles, research, breaches, news and events every week.

Bonus 33% off our NHI Course when you subscribe.

NHIMG Editorial Note
Reviewed and updated by the NHIMG editorial team on October 10, 2026.
NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org