Governance checklists confirm that a system was approved, but they do not control what happens when the AI is manipulating prompts, invoking tools, or moving through connected workflows. The risk appears at runtime, so security has to observe and constrain live behaviour, not just validate design-time intent.
Why governance checklists miss production AI behaviour
Governance checklists are mostly design-time artefacts: they tell you whether a model, workflow, or deployment was reviewed, approved, and documented. Production AI is different because the meaningful risk appears when the system starts interpreting prompts, chaining actions, calling tools, and traversing connected services. At that point, the control question is no longer “Was it sanctioned?” but “What can it do right now, and under what constraints?”
That distinction matters because runtime behaviour can change with context, data, retrieval results, tool availability, and user input. A checklist may confirm that a control exists in principle, yet still miss prompt injection, unintended tool invocation, over-broad action scope, or workflow escalation. For production systems, the security unit of analysis is live behaviour, not the approval record.
In practice, this is why teams that stop at governance paperwork often overestimate assurance. They may have policy statements, risk sign-off, and architecture diagrams, but they lack evidence that the AI can be observed, bounded, and interrupted while it operates. The control gap is operational: if the system can act across multiple tools or services, static approval does not show whether those actions remain safe under load, error, or adversarial prompting.
What runtime control has to add beyond checklist approval
Production AI security needs controls that are executable during execution, not just reviewable before release. That means constraining tool access, limiting which actions the system may take, validating inputs and outputs, and monitoring for unexpected transitions across workflows. The point is not to eliminate automation, but to make authority legible and bounded when the system is active.
This is where workload and agent identity become operationally important. If an AI workload can reach databases, APIs, ticketing systems, or deployment tools, its effective risk depends on the permissions attached to those pathways. A strong design can still fail if the runtime identity is too broad, poorly segmented, or reused across environments. AI Infrastructure Workload Identity Guide is useful here because it frames AI platforms as collections of jobs, registries, inference services, and compute that each need controlled access.
For AI systems that act on behalf of users or other services, you also need a clear model for delegation and retirement. The governance checklist may say “approved for production,” but runtime safety depends on whether the system’s authority is explicit, time-bounded, and revocable. Agentic AI Identity Guide helps explain why identity lifecycle and delegation are central once an AI can initiate actions rather than merely generate text.
At the control-layer level, the most reliable pattern is to pair approval with enforceable guardrails that match the actual integration surface. If the system uses APIs, tools, or connected workflows, you need policy enforcement, auditability, and a narrow operating envelope at those boundaries. That is why the SPIFFE workload identity specification is relevant: it shows how runtime identity and attestation can support stronger service-to-service trust than a checklist alone.
What fails when checklists replace live supervision
Checklist-driven governance fails in a few predictable ways. First, it assumes the primary risk is misdesign rather than misuse. Second, it treats the application as if it were static, even though production AI can change its execution path from one request to the next. Third, it often ignores the connected systems that become reachable once the model can invoke tools, retrieve data, or trigger downstream actions.
That creates a mismatch between approval and exposure. A system can be fully documented yet still be able to leak data, issue unintended requests, or cascade into business workflows that no reviewer anticipated. The more tools and shared services an AI can reach, the more the security question shifts from “Is the model allowed?” to “What exactly is it allowed to influence, and how quickly can that authority be reduced if behaviour changes?”
Governance also tends to underweight observability. If you cannot trace which prompt, retrieval result, tool call, or policy decision led to an action, then approval has limited operational value. Production AI needs evidence of execution paths, not just evidence of review. Without that, you cannot distinguish normal autonomy from emerging misuse, and you cannot learn whether a control failure is local, systemic, or recurring.
Risk and Threat Considerations
Production AI creates exposure when approved capabilities become broad enough to be abused at runtime. The main threat is not that the system exists, but that a prompt, retrieval result, or connected workflow can steer it into making harmful calls, crossing trust boundaries, or executing actions beyond the original intent.
Failure mechanism: A checklist validates the planned design, but does not prevent prompt injection, tool misuse, privilege overreach, or workflow chaining once the system is live and interacting with real inputs.
Impact: Attackers or accidental misuse can cause data exposure, unauthorized transactions, service disruption, or lateral movement across connected systems, even when the deployment was formally approved.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Agentic AI Top 10 addresses the attack and risk surface, while NIST SP 800-53 Rev 5 sets the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP Agentic AI Top 10 | ASI03 — Identity & Privilege Abuse | Production AI risk here centers on runtime authority and overreach. |
| ASI02 — Tool Misuse | The question is about tools and workflows being used unsafely at runtime. | |
| Recommendation — Constrain agent permissions and prevent privilege abuse during live execution. Restrict and monitor tool calls so the agent cannot misuse connected actions. | ||
| NIST SP 800-53 Rev 5 | AC-6 — Least Privilege | Runtime AI safety depends on minimizing what connected systems the workload can reach. |
| AU-6 — Audit Review, Analysis, and Reporting | The core gap is that static approval does not show what happened during execution. | |
| IA-9 — Service Identification and Authentication | Connected AI workflows rely on authenticated machine-to-machine access at runtime. | |
| Recommendation — Apply least privilege to every AI-facing service account and integration path. Review AI execution logs to detect unsafe prompts, tool use, and workflow escalation. Authenticate AI services and downstream tools with strong service identity controls. | ||
Practitioner Guidance
What to prioritise: Treat runtime authority as the primary control object. If the AI can invoke tools or affect records, start with the narrowest possible action scope and verify that every privileged path is observable and revocable.
What to verify: Check that approvals are backed by live controls, not just documentation. The practical test is whether you can explain, log, and constrain a real action taken by the system under adversarial or unexpected input.
Common mistake: Teams often over-trust model reviews and under-invest in execution controls. A signed-off design that cannot detect or stop unsafe tool use is still a weak production control.
Practitioner takeaway: Governance is necessary, but it is not a runtime security control; production AI becomes safe only when live behaviour is observable, bounded, and interruptible.
Related resources from NHI Mgmt Group
Deepen Your Knowledge
Free weekly newsletter
Subscribe to the NHI & AI Identity Journal
The latest on NHI and Agentic AI security – articles, research, breaches, news and events every week.
Bonus 33% off our NHI Course when you subscribe.
Reviewed and updated by the NHIMG editorial team on October 11, 2026.
NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org