The review model breaks because the skill can still drive runtime actions after approval. Prose instructions may trigger shell commands, file reads, remote fetches, or secret access that were never visible in a static review. The practical failure is treating the skill file as evidence of safety rather than as a control surface that must be governed during execution.
Why Documentation Review Fails for Agent Skills
Agent skills are not passive reference material. A skill file can define actions, tool calls, prompts, and branching logic that are only fully expressed when the agent executes them. Reviewing the prose alone misses the runtime behaviour that actually determines whether the skill can read files, invoke commands, fetch remote content, or reach secrets.
That means the review object is wrong if it treats the skill as a policy document. The correct object is an executable control surface: what it can do, what it can access, and what conditions cause those actions to fire. Static approval may still be useful, but it is only a snapshot of intent, not proof of safe operation.
Practically, this is the same category error seen when people review automation by reading configuration text instead of testing effective permissions. The meaningful security question is not whether the skill sounds acceptable, but whether its runtime path can expand privilege, cross trust boundaries, or trigger side effects that were invisible in the text review.
What Changes at Runtime
Runtime execution is where hidden risk appears. A skill may look harmless because it describes a simple workflow, yet it can still chain into shell execution, local file access, network retrieval, or downstream tool use once the agent starts interpreting it. The execution context, inherited credentials, and connected tools are what turn words into action.
That is why approval must account for effective authority, not just content. If the skill inherits a broad session, a connected connector, or a reused credential, then the resulting behaviour can exceed what the reviewer believed they were approving. The technical failure is not poor wording in the document, it is a mismatch between reviewed text and actual execution rights.
Skill review also needs to distinguish declared instructions from runtime dependencies. A skill may be safe in isolation yet unsafe when combined with other tools, broader context, or a permissive execution environment. The security boundary is therefore the agent runtime, not the file format.
Why a Static Approval Model Creates Blind Spots
Static review encourages false confidence because it focuses on visible prose and ignores latent capability. Once a skill is loaded, small edits in prompts, embedded commands, or tool mappings can change its behaviour without making the safety review obviously stale. That creates drift between the approved artifact and the operational reality.
It also weakens accountability. If the skill can act on behalf of a user or service, the critical question becomes who approved which permissions, for which scope, and with what revocation path. A document review process rarely answers those questions well unless it is tied to execution logging, approval state, and access review.
When a skill can reach files, APIs, or secrets, the risk is not theoretical. It can become a privilege boundary bypass, because the reviewer may bless the text while the runtime still inherits standing access that is much broader than the intended task.
Risk and Threat Considerations
Agent skills become risky when hidden instructions or inherited permissions allow them to cross the boundary from “described” behaviour to “effective” behaviour. The most material failures are secret exposure, unintended file access, and remote action through tools that were never visible in the review artifact.
Failure mechanism: The reviewer approves prose, but the execution environment interprets that prose with live credentials, connected tools, and broad operating scope. An attacker or careless author can exploit that gap by placing dangerous instructions where static review is unlikely to simulate the full runtime chain.
Impact: The agent can exfiltrate data, modify systems, or trigger actions that appear to be authorised because they originated from an approved skill. That turns the skill into a covert control plane, with consequences that range from data loss to delegated abuse of trusted access.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Agentic AI Top 10 and OWASP Agentic Skills Top 10 address the attack and risk surface, while NIST SP 800-53 Rev 5, OWASP ASVS and NIST Zero Trust (SP 800-207) set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP Agentic AI Top 10 | ASI03 — Identity & Privilege Abuse | Agent skills can exercise runtime authority beyond static review. |
| ASI02 — Tool Misuse | The core failure is hidden tool invocation from apparently safe skill text. | |
| ASI05 — Unexpected Code Execution | Skills may trigger shell commands or other runtime execution paths. | |
| Recommendation — Enforce per-action authorization and limit inherited privilege for skill execution. Constrain which tools a skill may invoke and validate each call path. Sandbox execution paths and block unreviewed command execution. | ||
| OWASP Agentic Skills Top 10 | AST10 — Agentic Skills Top 10 | The question is directly about the security of the agent skill layer. |
| Recommendation — Review skill registries and execution semantics, not prose alone. | ||
| NIST SP 800-53 Rev 5 | AC-6 — Least Privilege | Skills should not inherit broad authority beyond their intended task. |
| AU-2 — Event Logging | Runtime action visibility is needed to distinguish safe text from unsafe behaviour. | |
| Recommendation — Restrict each skill to the minimum effective permissions needed. Log skill-triggered actions and retain evidence for review and incident response. | ||
| OWASP ASVS | V15 — Secure Coding and Architecture | Executable agent skills need architecture-level review of side effects and boundaries. |
| Recommendation — Design agent skill execution so control flow and side effects are explicit and bounded. | ||
| NIST Zero Trust (SP 800-207) | Zero Trust Architecture | The answer depends on verifying runtime actions rather than trusting approved text. |
| Recommendation — Verify each action at runtime and avoid assuming approved skills are safe by default. | ||
Practitioner Guidance
What to verify: Treat every approved skill as an executable object and confirm its effective permissions, not just its text. The review should answer what it can touch, what it can invoke, and what credential or session context it inherits at runtime.
Common mistake: Do not rely on manual reading as the primary safety gate when the skill can branch into tools or commands. A safe-looking description is not evidence that the runtime path is bounded, observable, or revocable.
What good looks like: The skill is approved only with explicit scope, tested execution paths, telemetry on actions taken, and a clear revocation or disablement path. If the runtime cannot be observed and constrained, the approval is incomplete.
Practitioner takeaway: Review the skill the way you would review code with production access, because the real risk is not the prose itself, but the authority it can exercise once executed.
Related resources from NHI Mgmt Group
Deepen Your Knowledge
Free weekly newsletter
Subscribe to the NHI & AI Identity Journal
The latest on NHI and Agentic AI security – articles, research, breaches, news and events every week.
Bonus 33% off our NHI Course when you subscribe.
Reviewed and updated by the NHIMG editorial team on October 11, 2026.
NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org