TL;DR: OWASP’s Agentic Skills Top 10 (AST10) maps ten risks to 206,435 analysed skill files, showing that skills can execute shell commands, read SSH keys, and reach the internet, with 23.7% reading sensitive local stores and 4.7% using any sandboxing, according to Capsule. Static review is not enough when the payload is prose and trust becomes standing.
At a glance
What this is: OWASP’s Agentic Skills Top 10 formalises ten risks in agent skills, backed by Capsule’s analysis of 206,435 files showing that skills already behave like executable software with real access paths.
Why it matters: IAM and NHI teams need to treat agent skills as governed execution surfaces because approval alone does not control what a skill can read, run, or reach at runtime.
By the numbers:
- 23.7% of skills read sensitive local stores such as ~/.ssh, cloud credentials, the macOS Keychain, or browser cookies.
- 4.7% of skills use any sandboxing or containment.
- 3.6% of skills require a human approval step.
Context
Agent skills are reusable behaviors that agents install and obey, which makes them closer to software modules than to documentation. The risk is not only what a skill says, but what it can cause an agent to execute, fetch, or expose once trusted.
This article focuses on the governance gap created when skill review, metadata, and approval are treated as sufficient controls. For agentic AI programmes, the key question is whether the skill itself has runtime reach into shell commands, files, credentials, and outbound communications.
Capsule’s analysis is framed around a large sample of public skill files, but the underlying issue is broader than one ecosystem. Any programme allowing agents to install or follow external instructions inherits the same control problem: static review cannot bound runtime behavior.
Key questions
Q: What breaks when agent skills are reviewed like documentation instead of executable software?
A: The review model breaks because the skill can still drive runtime actions after approval. Prose instructions may trigger shell commands, file reads, remote fetches, or secret access that were never visible in a static review. The practical failure is treating the skill file as evidence of safety rather than as a control surface that must be governed during execution.
Q: Why do agent skills create more risk when they can reach SSH keys or cloud credentials?
A: Because the skill inherits the authority of the secrets it can touch. Once a skill can read SSH material, cookies, keychains, or cloud credentials, a single malicious instruction can move from local execution into broader identity compromise. That expands blast radius, lateral movement potential, and persistence options far beyond the original install event.
Q: How should teams decide whether to trust a skill that points to remote instructions or dependencies?
A: They should not rely on the reference alone. Remote instructions and dependency links are mutable trust inputs, so the key decision is whether the skill’s behaviour is pinned, monitored, and revocable after install. If the content can change without reapproval, the trust decision is incomplete.
Q: What runtime controls matter most for agentic skills that can execute shell commands?
A: The most useful controls are the ones that can stop the action at execution time, not just at intake. That means policy enforcement for shell commands, outbound connections, and data writes, plus visibility into which skills can touch which identities and secrets. If you cannot interrupt the action, you do not actually control the skill.
Technical breakdown
Why agent skills behave like executable code
An agent skill is usually packaged as prose plus instructions, but the agent interprets that prose as operational authority. If the skill can instruct shell execution, file reads, or remote fetches, the real security boundary is not the file format. It is the set of actions the agent is willing to perform after reading it. That is why skills that look like markdown can still create code execution, data access, and outbound communication paths. The control problem is not documentation hygiene. It is governance over instructions that become runtime behaviour once installed.
Practical implication: govern skills as executable inputs, not as static content assets.
Why metadata and approval do not bound skill risk
Metadata can suggest safety, but it cannot prove safe execution. A skill listing can be misleading, a dependency can drift, and an approval step can only certify the state visible at that moment. If the skill later pulls remote content or changes behaviour through external instructions, the trust decision has already been converted into standing exposure. The same issue appears when a skill is reviewed once and then reused across environments without rescanning. The governance failure is treating a snapshot as a lifecycle control.
Practical implication: pair intake review with ongoing runtime monitoring and version control.
How weak isolation and over-privilege amplify agent reach
Skills become dangerous when they inherit broad local authority. Access to SSH material, cloud credentials, browser cookies, or keychains turns a single instruction file into a bridge across identity domains. Weak isolation makes that worse because the agent can act with fewer barriers between the skill, the host, and external services. In practice, over-privilege and low containment are multiplicative: each one increases the blast radius of the other. For agentic AI, least privilege has to apply to the skill’s effective reach, not just to the host account that launched it.
Practical implication: reduce the skill’s reachable surfaces before you rely on human approval.
Threat narrative
Attacker objective: The attacker wants the agent to execute untrusted instructions with enough authority to steal credentials, exfiltrate data, or establish downstream control.
- Entry occurs when a malicious or compromised skill is installed through a registry, repository, or remote reference that the agent is configured to trust.
- Credential or command access follows when the skill directs the agent to read local secrets, execute shell commands, or fetch attacker-controlled instructions.
- Impact occurs when the agent carries out the skill’s instructions with real permissions, enabling theft, persistence, or external communication from the host environment.
Breaches seen in the wild
- Carbonato botnet 2026: A botnet puts a Hermes AI agent on exposed Docker hosts to steal AI API keys, SSH keys and tokens, powering the operators' own LLM gateway.
- GhostAction returns 2026: GhostAction's fake security workflow hit 772 repos in Aug-Sept 2026 via stolen GitHub credentials, targeting 2,577 secrets; only 16% cleaned.
Read and download The State of NHI & AI Agent Breach Report 2026, covering 200+ breaches impacting Non-Human Identities including AI Agents.
NHI Mgmt Group analysis
Agent skills collapse the old documentation-versus-code distinction: once a skill can instruct shell execution, credential reads, or remote fetches, it is no longer safe to govern as text. The article’s data shows that the ecosystem already treats prose as runtime authority. Practitioners should assume every installable skill is an executable governance object, not a note attached to the agent.
Static review creates a false sense of control: approval, metadata, and snapshot scanning all presuppose that trust can be decided before execution and remain valid afterwards. Skills break that assumption because remote references, dependency drift, and content updates can change behaviour after review. The result is trust that is standing, not transactional, which means governance must follow the skill into runtime.
Over-privileged skills are the fastest path to identity abuse: when a skill can reach SSH keys, browser cookies, or cloud credentials, it inherits the same blast radius as the identities those secrets unlock. That makes the skill layer part of NHI governance, not a separate AI problem. Security teams need to map skill permissions to the identities and secrets they can touch, then treat that mapping as a control boundary.
Runtime governance gap: the article shows that the field has invested in lists, labels, and reviews faster than it has invested in enforcement. That gap is now the category’s defining failure mode. The practitioner conclusion is straightforward: if a skill can execute, then it needs runtime controls that can still stop it after intake approval has already succeeded.
Agentic supply chain risk is becoming content-driven: the payload may live in prose, metadata, dependency references, or registry reputation signals rather than in code. That broadens the threat surface beyond scanners built for source code and packages. Teams should treat agent skill ecosystems as supply chains with identity consequences, not as a niche prompt-injection problem.
What this signals
Runtime governance, not intake review, is the decisive control boundary for agent skills: skills can change behavior through prose, remote references, and dependency drift after they have already passed approval. That means the security programme has to observe what the skill does in session, not just what it looked like when submitted.
Agentic supply chain risk now includes identity reach: when a skill can touch SSH material, browser cookies, or cloud credentials, it becomes part of the NHI attack surface. Practitioners should track the secrets each skill can reach and tie that inventory to enforcement, not to policy documents.
Static scanners miss the payload layer when the payload is natural language: the article shows that harmful instructions can live in prose, which makes code-only inspection insufficient. Teams that rely on source scanning alone will keep missing the control plane where agent behavior is actually decided.
For practitioners
- Inventory installed agent skills Build a register of every skill your agents can install or follow, including prose-only instructions and remote references, so you can govern the actual attack surface.
- Gate credential and network reach Block skills from reading SSH keys, browser cookies, keychains, or cloud credentials unless that reach is explicitly approved and continuously monitored.
- Pin skill versions and dependency references Treat remote docs and dependency pointers as mutable inputs and lock them to verified versions before they can change agent behaviour.
- Add runtime controls for skill execution Use policy enforcement that can interrupt unsafe shell commands, outbound connections, and database writes after a skill has passed intake review.
Key takeaways
- Agent skills have crossed the line from documentation into executable behavior, which changes how identity teams must govern them.
- The article shows that over-privilege, weak isolation, and mutable remote references are already common in skill ecosystems.
- Runtime enforcement is the control that changes outcomes, because approval and static review do not stop harmful action once execution begins.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Agentic AI Top 10 and OWASP Non-Human Identity Top 10 define the specific risk controls and attack patterns relevant to this term.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP Agentic AI Top 10 | ASI02 — Tool Misuse | Skills direct agents to misuse tools, files, and shell access across the workflow. |
| ASI03 — Identity & Privilege Abuse | The article centers on skills inheriting access to SSH keys, cookies, and cloud credentials. | |
| Recommendation — Map skill permissions to ASI02 and restrict which tools each skill can invoke at runtime. Apply ASI03 controls to prevent skills from expanding into identities and secrets they were not meant to reach. | ||
| OWASP Non-Human Identity Top 10 | NHI-02 — Secret Leakage | The article shows skills reading SSH keys, credentials, and browser cookies. |
| NHI-05 — Overprivileged NHI | Skills frequently inherit more authority than their task needs, creating excess blast radius. | |
| NHI-07 — Long-Lived Secrets | Remote references and unpinned dependencies keep skill trust standing long after review. | |
| Recommendation — Use NHI-02 to stop skills from exposing secrets through local reads, remote fetches, or command execution. Apply NHI-05 to reduce the reachable files, commands, and network paths available to each skill. Treat NHI-07 as a trigger to shorten credential and reference lifetimes tied to skills. | ||
Key terms
- Agent Skill: A reusable package of task-specific knowledge and procedures that an autonomous agent can load when needed. In practice, it separates general awareness from operational detail, which makes enterprise context easier to govern than a single oversized prompt.
- Runtime Governance: Runtime governance is the set of controls that verify what a system or agent is actually doing after deployment. It combines monitoring, authorization checks, and access validation so teams can detect drift, misuse, or excessive privilege in motion rather than assuming build-time policy still holds.
- Skill Drift: The divergence that occurs when the same AI workflow artefact exists in multiple copies and no longer behaves consistently across tools or teams. It creates hidden policy variance, weakens auditability, and makes it hard to know which version is authoritative.
- Agentic Supply Chain: The collection of models, tools, plugins, prompts, memory stores, and middleware that an AI agent depends on to operate. Weaknesses in this chain can introduce hidden instructions, poisoned context, or exposed secrets, so security teams need inventory, trust validation, and isolation controls across the entire path.
Deepen your knowledge
NHI governance, agentic AI identity, and machine identity lifecycle are core topics in our NHI Foundation Level course, the industry's only accredited NHI security programme. If you are building or maturing an IAM programme, it is worth exploring.
Published by the NHIMG editorial team on October 11, 2026.
NHI Mgmt Group, the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org