TL;DR: AI agent approvals fail when teams rubber-stamp declared permissions instead of proving observed behaviour, baseline coverage, and response readiness, according to ARMO. The decisive issue is not whether an agent exists, but whether its runtime evidence closes the gap between intended access and actual tool use.
At a glance
What this is: This is an operational checklist for approving AI agents for production, and its central finding is that go-live decisions need evidence gates, not framework posters.
Why it matters: IAM, PAM, and NHI teams need a concrete approval artefact because agent identity risk changes with autonomy tier, observed behaviour, and dormant access surface.
By the numbers:
- Only 44% of developers are reported to follow security best practices for secrets management, exposing a significant developer behaviour gap.
👉 Read ARMO's checklist for AI agent production approval gates
Context
AI agent production approval has become an identity governance problem, not just an application release problem. The core issue is that declared access, observed behaviour, and actual blast radius rarely align closely enough for a safe go-live decision, especially when tool-calling agents can change their effective reach during runtime.
A CISO checklist matters because current AI security programmes often describe what should be built, but not what evidence should be presented at approval time. For NHI and agentic AI governance, the missing control is usually not a policy statement. It is a runtime-derived approval record that proves the agent’s real access, behaviour baseline, and response readiness before production.
ARMO frames the question as an operational go-live review for AI agents. That starting point is typical of teams that have begun runtime observability work but still lack a consistent evidence standard for approvals.
Key questions
Q: How should security teams evaluate AI agent trust before production use?
A: Security teams should evaluate AI agent trust by combining identity posture, intended access, delegation paths, and governance metadata in one approval decision. The key question is not whether the agent is useful, but whether its runtime behaviour can be bounded before it reaches sensitive systems. Registries can support that decision, but they do not replace policy enforcement.
Q: Why do AI agents create access problems that human approval processes do not solve well?
A: AI agents can inherit credentials, chain actions, and execute at machine speed, while human approval models assume slower, request-based behaviour. That mismatch creates delays, workarounds, and eventually shadow AI access. The control problem is not whether to approve access, but how to do it fast enough without losing visibility or governance.
Q: What breaks when declared permissions do not match observed AI agent behaviour?
A: The entire approval decision breaks because the team is certifying access it has not validated. A mismatch means dormant tools, unseen data paths, or hidden write access may already exist. Once that delta is accepted, every downstream control inherits false confidence, including detection tuning and incident response planning.
Q: How do organisations know when an approved AI agent needs re-review?
A: Re-review is needed when the agent’s prompt, model, tools, or reachable data changes enough to alter its behaviour baseline. Security teams should also re-check after new integrations, scope expansion, or unexpected access patterns. In practice, any drift from the approved runtime profile should trigger a fresh decision.
Technical breakdown
AI-BOM inventory and autonomy tiering
An AI bill of materials, or AI-BOM, is the canonical inventory of what an agent can load, call, and touch at runtime. For approval purposes, that inventory has to include models, RAG sources, MCP-connected tools, APIs, owner, reviewer, and autonomy tier. The important distinction is between intended use and reachable capability. If a supposedly read-only agent retains write credentials or dormant tool access, the inventory must reflect the reachable surface, not the project description.
Practical implication: require runtime-derived inventory entries before any production approval, and classify the agent upward when tool reach is ambiguous.
Behavioural baselines and declared-vs-observed access
A behavioural baseline is a measured record of how an agent actually behaves across representative traffic, including tool sequences, destinations, data access, and process activity. The checklist’s most useful control is the declared-vs-observed comparison, which exposes when an agent is authorised for more than it truly needs or is already touching more than engineering admitted. For autonomous or tool-rich agents, the baseline is a probability distribution, not a single deterministic path, so short observation windows undercount exposure.
Practical implication: approve only when the observed tool set and data paths match the declared scope, and treat dormant permissions as latent blast radius.
AI-native detection and re-approval triggers
AI-native detection is not generic container alerting with an AI label attached. It has to identify prompt injection, tool misuse, data exfiltration through legitimate outputs, and changes in the agent’s normal action graph. Re-approval triggers matter because agent behaviour can drift after approval through prompt changes, tool additions, or model updates. In identity terms, approval is a snapshot, while runtime control has to be continuous.
Practical implication: tie detection coverage and drift-triggered re-approval into the release process so approval does not become a one-time paper exercise.
Threat narrative
Attacker objective: The attacker objective is to turn an approved AI agent into a governed-but-overreaching execution path that reaches data or systems beyond the intended boundary.
- Entry begins when a user or upstream content reaches the agent prompt and induces a tool-calling path the security team did not intend.
- Escalation follows when the agent’s declared permissions and observed behaviour diverge, allowing dormant tools or broader data paths to become reachable.
- Impact is realised through unauthorised data access, write actions, or chained agent activity that expands blast radius beyond the approved use case.
Breaches seen in the wild
- Meta AI Instagram Account Takeover — 20,225 Instagram accounts hijacked via compromised Meta AI support chatbot with overprivileged access.
- Replit AI Tool Database Deletion — Replit vibe coding AI assistant deletes live production database and creates 4,000 fake user records.
Read our 52 NHI Breaches Analysis report for a comprehensive view of breaches impacting Non-Human Identities including AI Agents.
NHI Mgmt Group analysis
Declared access is not an approval control when observed behaviour diverges: The checklist’s strongest insight is that approval decisions fail when teams certify intended permissions instead of runtime reach. If an agent is declared for 47 APIs but only shows 3 in observation, the missing 44 are not theoretical. They are ungoverned surface area that security has not actually measured. The practitioner conclusion is that go-live approval must be evidence-led, not narrative-led.
Autonomy tiering collapses the idea of one-size-fits-all AI governance: A read-only retrieval assistant and a multi-agent workflow system do not belong on the same approval standard. The more independent the runtime decisions, the less useful static documentation becomes as a control. That means identity programmes have to stop treating AI agents as a single category and start mapping approval depth to actual decision authority.
Identity blast radius is the real approval variable: The dangerous question is not whether the agent is modern or well instrumented. It is how far its privilege can travel when prompt content, tool chaining, or upstream data change its behaviour. This is the same structural problem NHI teams face with service accounts that are provisioned broadly and reviewed too late. The practitioner conclusion is to approve the reachable radius, not the intended role.
Runtime evidence should replace paper assurance in every go-live meeting: The article is right to push a concrete gate checklist because many AI security frameworks stop at program design. That leaves CISOs with no artefact to sign in the room where production decisions get made. The field needs operational approval standards that combine inventory, baseline, detection, and re-approval logic into one decision record.
From our research:
- The average estimated time to remediate a leaked secret is 27 days, despite 75% of organisations expressing strong confidence in their secrets management capabilities, according to The State of Secrets in AppSec.
- Organisations maintain an average of 6 distinct secrets manager instances, which fragments control and slows consistent governance across workloads.
- That same gap between confidence and reality is why teams should read the Ultimate Guide to NHIs , 2025 Outlook and Predictions alongside approval checklists when agent access is in scope.
What this signals
AI agent approval is becoming a governance function: teams that treat go-live as a product checklist will keep missing the identity evidence that actually matters. Runtime inventory, baseline drift, and declared-vs-observed access need to sit inside the release decision, not outside it. That is especially true once agents begin to touch internal tools, customer data, or downstream workflows.
Runtime evidence changes the operating model for both NHI and agentic AI: the same control logic that exposes dormant secret exposure in NHI estates also exposes hidden reach in AI agents. If your programme still relies on static entitlements and periodic review alone, the approval model is already behind the behaviour model.
The practical shift is toward continuous re-approval, not one-time sign-off. For identity teams, that means aligning agent approvals with OWASP NHI Top 10 style threat thinking and with NIST AI Risk Management Framework concepts where autonomous behaviour changes the control surface.
For practitioners
- Build a runtime AI-BOM before production review Record the agent’s loaded models, RAG sources, MCP tools, reachable APIs, named owner, and autonomy tier from runtime discovery rather than manual self-reporting.
- Compare declared permissions to observed behaviour Approve only when observed tool use, network destinations, and data paths align with the declared access set, and treat any delta as unresolved blast radius.
- Set tier-based evidence thresholds Use a stricter observation window and coverage requirement for higher-autonomy agents, because one week of logs is not enough for a workflow agent with write access.
- Tie detection to agent-specific failure modes Require rules that flag prompt injection, tool misuse, and data exfiltration through legitimate outputs, not just generic network or container alerts.
- Make re-approval automatic on drift signals Trigger review when prompts change, tools are added, models are swapped, or the agent starts reaching new data sets outside its original baseline.
Key takeaways
- AI agent go-live decisions fail when teams approve intent instead of runtime evidence.
- Declared permissions, observed behaviour, and blast radius need to match before production sign-off.
- The approval model now belongs inside identity governance, because drift turns a signed-off agent into an unmanaged access path.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Agentic AI Top 10 and OWASP Non-Human Identity Top 10 address the attack and risk surface, while NIST CSF 2.0, NIST AI RMF and NIST Zero Trust (SP 800-207) set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP Agentic AI Top 10 | The article centres on production approval risk for agentic systems and tool misuse. | |
| OWASP Non-Human Identity Top 10 | NHI-03 | Declared-vs-observed access and secret exposure are core non-human identity controls here. |
| NIST CSF 2.0 | PR.AC-4 | The checklist is fundamentally about access permissions and approval evidence. |
| NIST AI RMF | GOVERN | Production approval depends on accountability, oversight, and governance for AI systems. |
| NIST Zero Trust (SP 800-207) | The checklist uses continuous verification logic consistent with zero trust for runtime access. |
Treat each agent tool call as a fresh trust decision and limit access to the minimum required per session.
Key terms
- AI-BOM: An AI bill of materials is a structured inventory of the components that define an AI agent, including the model, prompt, tools, retrieval sources, and dependencies. In practice, it is the evidence base for review, change control, and risk assessment when the agent evolves after deployment.
- Declared-vs-observed access: Declared-vs-observed access is the comparison between what an agent is supposed to touch and what runtime evidence shows it actually touches. The gap matters because dormant permissions and hidden data paths often create more risk than the design documents suggest.
- Behavior Baseline: A record of normal activity for a non-human identity, including typical consumers, resources, and actions over time. Baselines help security teams detect when an identity is being used in an unusual way and provide the context needed to enforce least privilege safely in dynamic environments.
- Re-approval trigger: A re-approval trigger is any change that invalidates the assumptions behind a prior production sign-off. In agent governance, that usually means prompt changes, tool additions, model swaps, or new data access that alters the original risk profile.
What's in the full article
ARMO's full blog covers the operational detail this post intentionally leaves for the source:
- The full seven-gate approval checklist with concrete evidence standards for each gate.
- Autonomy-tier thresholds that show how evidence bars change between read-only agents and multi-agent workflows.
- The one-page approval artifact template used to turn a go-live discussion into a signed decision record.
- Practical examples of declared-vs-observed mismatches and what security reviewers should look for in the console.
👉 ARMO's full post covers the seven gates, evidence thresholds, and the approval artefact template.
Deepen your knowledge
NHI governance, agentic AI identity, and machine identity lifecycle are core topics in our NHI Foundation Level course, the industry's only accredited NHI security programme. If you are building or maturing an identity security programme, it is worth exploring.
Published by the NHIMG editorial team on August 20, 2026.
NHI Mgmt Group — the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org