Treat the run, not the agent label, as the governance unit. When work can pause, resume, and change context over time, access approval, evidence capture, and rollback need to attach to the execution instance and its version. Otherwise, the team cannot prove what authority existed at each stage of the task.
Why run-level governance matters for long-lived agent work
Security teams should govern the execution run as the accountable unit because the run is what accumulates authority, context, side effects, and evidence over time. An agent label is too coarse once work can pause, resume, branch, or pick up new inputs hours later. The question is not only who started the task, but what authority existed at each point in the run.
That distinction becomes important when approvals age out, inputs change, or the task crosses control boundaries. The same named agent can behave differently across resumptions if its session, delegated permissions, or tool access has changed. Governance therefore needs a run identifier, a versioned policy snapshot, and a way to tie every meaningful action back to the exact execution state that produced it.
Long-running work also creates a record-keeping problem: if evidence is attached only to the agent identity, teams lose the ability to reconstruct what was authorised, observed, and approved during each stage. A run-centric model preserves auditability because it links decisions to the specific instance that executed them, rather than to a reusable label that may outlive the original context.
What has to be versioned and why
The minimum governance set is the task definition, the approval context, the permissions in force, and the evidence trail. Each of those can change independently during a long task, so the control model should treat them as versioned artefacts, not static metadata. If the run resumes after interruption, the resumed segment should inherit only what is still valid, not whatever authority happened to exist when the task first began.
This is especially important when the execution can interact with external systems, call tools, or act on behalf of a user. A resumed run may need a fresh decision on access scope, current business justification, and whether the original approval still covers the new step. The safest operational pattern is to re-evaluate permission at the point of action, while preserving the earlier state as evidence rather than assuming it remains valid.
Versioning also makes rollback more precise. When a task must be stopped or reversed, teams need to know which actions belong to which run version so they can revoke the right access, isolate the right outputs, and distinguish completed work from partially executed work. That is much harder if multiple resumptions are blended into one undifferentiated agent history.
How to design governance for pause, resume, and recovery
Governance should be built around state transitions: start, pause, resume, escalate, and stop. Each transition should create or refresh an auditable checkpoint so the system can prove what was allowed before and after the break. For hours-long work, the checkpoint matters as much as the approval itself because context decay is inevitable.
Where the work can continue after interruption, the control question becomes whether the resumed run is still the same authorised activity. If the answer depends on timing, external data, or changing scope, then the system needs a re-approval rule or a bounded continuation window. That reduces the chance that stale authority silently carries forward into a new context.
Teams should also define what survives a pause. Some state, such as task progress or non-sensitive intermediate outputs, may be safe to retain, while other state, such as short-lived credentials or high-risk tool permissions, should expire with the checkpoint. The practical goal is to make resumption reliable without letting a stale execution inherit more power than the current moment justifies.
Risk and Threat Considerations
Long-running and resumable agent work increases the chance of stale authority, ambiguous accountability, and hidden scope creep. If access is tied to a generic agent identity instead of the execution run, a resumed task can continue under assumptions that no longer hold, especially after approval expiry, context change, or partial compromise.
Failure mechanism: The control failure is a mismatch between the authority that existed when the task began and the authority actually in force when the next action occurs, which can let a resumed run exceed its intended scope or obscure which step caused a harmful change.
Impact: Teams may be unable to prove who authorised a step, what permissions were active, or how far the run progressed before interruption, which weakens auditability, rollback, and incident containment.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Agentic AI Top 10 addresses the attack and risk surface, while NIST SP 800-53 Rev 5 and NIST Zero Trust (SP 800-207) set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP Agentic AI Top 10 | ASI03 — Identity & Privilege Abuse | Run-level governance prevents authority from drifting across resumed agent work. |
| ASI08 — Cascading Failures | Interrupted agent runs can propagate bad state into later steps and recoveries. | |
| ASI10 — Rogue Agents | Unchecked resumption can let an execution continue beyond its approved intent. | |
| Recommendation — Bind approvals to each execution run and re-check privilege before every resumed action. Checkpoint run state so recovery cannot reuse stale context or permissions. Require fresh authorization when a resumed run changes scope or timing. | ||
| NIST SP 800-53 Rev 5 | AU-3 — Content of Audit Records | The question depends on proving what authority and actions existed during each run stage. |
| AU-12 — Audit Record Generation | Long-running agent work needs event records that survive pauses and resumptions. | |
| AC-6 — Least Privilege | Resumed work should not inherit more access than the current step requires. | |
| Recommendation — Record run version, approvals, and action context in each audit event. Generate auditable checkpoints at start, pause, resume, and stop. Scope permissions to the current run segment and revoke surplus access immediately. | ||
| NIST Zero Trust (SP 800-207) | PA — Policy Engine / Decision Point | Run-time decisions should be re-evaluated at each resumed action, not assumed from startup. |
| Recommendation — Enforce policy decisions at action time for every resumed task step. | ||
Practitioner Guidance
What to prioritise: Define the run as the smallest unit that can independently receive approval, execute actions, and produce evidence. If the task can pause and resume, require the resumed segment to carry its own checkpointed context and policy state, even when the same agent continues the work.
What to verify: Make sure logs, approvals, and rollback records bind to the execution instance version, not just the agent name. If a reviewer cannot reconstruct the authorised scope at each stage, the governance model is too coarse for the task profile.
Practitioner takeaway: Durable agent work is governed safely only when authority is time-bounded, versioned, and attached to the specific run that exercised it, because continuity of identity does not guarantee continuity of permission.
Related resources from NHI Mgmt Group
Deepen Your Knowledge
Free weekly newsletter
Subscribe to the NHI & AI Identity Journal
The latest on NHI and Agentic AI security – articles, research, breaches, news and events every week.
Bonus 33% off our NHI Course when you subscribe.
Reviewed and updated by the NHIMG editorial team on October 11, 2026.
NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org