They miss behaviors that appear only during execution, such as unauthorized network calls, file writes, credential access, or data exfiltration. A skill can look clean in review and still act maliciously once it runs with the agent’s inherited permissions. That is why source review alone is not a trust decision for agent skills.
Why source-only scanners miss the real failure mode in agent skills
Source inspection can tell you whether a skill contains obvious bad code, but it cannot prove what the skill will do once it is invoked inside a live agent runtime. The meaningful failure mode is execution-time behavior, especially when the agent inherits permissions, tokens, network reach, or filesystem access that are not obvious from static review alone.
The key issue is that agent skills are often only one step in a larger trust chain. A clean-looking implementation can still call external services, write files, read secrets, or trigger downstream actions only after the agent supplies context and authority at runtime. That is why the security question is not just “what does the code say?” but “what can it do in the environment where it runs?”
In practice, that distinction matters for skills that rely on delegated permissions and tool access. A scanner that stops at source text will miss runtime-only abuse patterns such as unauthorized network calls, hidden data movement, and actions that appear harmless until the agent’s inherited privileges are attached. For that reason, AI agent authorisation guidance has to focus on per-action decisions, not just code review.
What execution-time behavior changes the answer
The important behaviors are the ones that emerge after the skill is loaded, chained, or granted access. That includes outbound requests, file system writes, credential use, hidden prompts to other tools, and attempts to move data out of the expected boundary. A source scan may show none of that because the dangerous branch only executes under particular inputs or permissions.
This is why trust decisions for agent skills need to include runtime context, not just artifact inspection. If a skill can inherit a session, reuse an existing token, or reach a tool with broader access than the code appears to need, the real risk is privilege amplification at execution time. Practical controls therefore need observability over agent actions, not only code provenance. AI agent observability, audit and incident response guidance is valuable here because it centers attribution, logging, and kill-switch decisions on what the agent actually did.
That same logic applies to skills that look benign in review but become destructive once bound to production permissions. If execution can touch sensitive systems, the skill’s safety depends on scope, context, and enforcement, not on the cleanliness of the source tree. A runtime policy that constrains what the agent can call is more meaningful than a static “approved” label on the skill package.
Why source review alone is the wrong trust boundary
Source review is still useful, but only as one input into a broader control set. It can help catch obvious malicious payloads, unsafe dependencies, or suspicious hard-coded endpoints. It cannot reliably detect logic that is intentionally deferred until execution, triggered by a prompt, or activated through tool chaining inside the agent.
That is especially true for skills in agentic systems, where the real exposure comes from delegated authority. The risk is not simply that the code exists, but that the agent can execute it with inherited access to accounts, APIs, files, or network paths the author never visibly declared. Zero trust for AI agents is the right mental model because it treats each request as something to verify, not something to trust because the code was inspected once.
For the same reason, runtime policy and least privilege matter more than reviewer confidence. If a skill can perform sensitive actions, then the control objective is to limit blast radius, record what happened, and make abuse visible quickly. That is also why the security boundary needs to sit around execution and authorization, not just the source repository.
Risk and Threat Considerations
The core risk is that a skill can pass static review and still behave maliciously only when it is executed with live permissions. That creates a false sense of safety, especially when the skill can reach secrets, internal APIs, or filesystems that are unavailable in the source artifact itself.
Failure mechanism: The scanner evaluates code in isolation, while the real attack or abuse path appears only after runtime context, agent delegation, and inherited permissions are attached. That gap lets hidden network access, exfiltration, or destructive actions evade source-only checks.
Impact: An apparently safe skill can still leak data, modify systems, or misuse credentials under the agent’s authority, which means approval based only on source review can produce unauthorized access and high-blast-radius compromise.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Agentic AI Top 10 and MITRE ATT&CK address the attack and risk surface, while NIST SP 800-53 Rev 5 sets the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP Agentic AI Top 10 | ASI03 — Identity & Privilege Abuse | Agent skills can abuse inherited authority only at runtime. |
| ASI02 — Tool Misuse | The failure mode is malicious tool use that static source review may miss. | |
| ASI10 — Rogue Agents | A skill can behave safely in review and still act adversarially when executed. | |
| Recommendation — Enforce per-action authorization and restrict inherited privileges before enabling agent skills. Validate tool calls at runtime and block unsafe tool invocation paths. Monitor agent behavior continuously and isolate untrusted agent actions. | ||
| MITRE ATT&CK | T1020 — Data Exfiltration | The question explicitly includes execution-only exfiltration as a missed behavior. |
| Recommendation — Hunt for outbound transfer patterns and alert on unexpected data movement. | ||
| NIST SP 800-53 Rev 5 | AC-6 — Least Privilege | Runtime privilege scope determines whether a reviewed skill can do harm. |
| Recommendation — Limit each skill to the minimum permissions needed for the task. | ||
Practitioner Guidance
What to verify: Treat source review as necessary but insufficient. Verify what the skill can actually reach at runtime, which permissions it inherits, and whether the execution path is constrained by per-action policy rather than package approval alone.
Decision rule: If a skill can invoke tools, reach the network, or touch sensitive data after deployment, require runtime controls, logging, and revocation paths before trusting it in production. If the skill’s behavior depends on agent context, test it in the same authorization environment where it will run.
Common mistake: Teams often equate “no suspicious source code” with “safe to enable.” For agent skills, the better question is whether the runtime can force the same behavior limits that the reviewer assumed.
Practitioner takeaway: Static inspection is a code-quality check, not a trust decision; the security decision belongs to the combination of runtime permissions, observability, and enforced action boundaries.