TL;DR: Unosecur says Kimi K3 exposed that the real agent security failure sat in the harness, where authentication bypass, path traversal, MCP permission gaps, and auto-approval behaviours could trigger authorised actions without a jailbreak. The boundary problem is not prompt refusal but who can turn model output into external authority.
At a glance
What this is: This analysis argues that Kimi K3’s launch exposed an agent identity problem in the harness, not just a jailbreak problem, because confirmed auth and filesystem flaws sat beside auto-approval and MCP permission behaviours.
Why it matters: IAM and security teams need to separate model refusal from tool authority, because agent security fails when model output can still become an authorised action path.
👉 Read Unosecur's analysis of Kimi K3, MCP harness controls, and agent identity risk
Context
Kimi K3’s launch is best understood as an agent identity and authorisation problem, not simply a prompt-safety problem. The issue is whether a model, its harness, and its tools are aligned so that only intended actions become executable.
Moonshot AI’s Kimi Code bundled MCP servers, skills, plugins, hooks, and sub-agents into one execution environment, which means authority was distributed across several control points at once. When authentication checks, filesystem access, and auto-approval sit in that same chain, the governance question becomes where the real enforcement boundary lives.
That is why the article separates jailbreaks from harness flaws. A jailbreak tests model refusal behaviour, but an authenticated tool path tests whether the agent can still perform actions it should never have been able to reach in the first place.
Key questions
Q: What breaks when an agent harness treats session approval as standing access?
A: Session approval stops being a narrow control and becomes reusable authority for later tool calls. That creates a hidden privilege window where the agent can act on changing context without fresh intent checks. The practical failure is not only overreach, but the loss of a clean approval boundary for sensitive actions.
Q: Why do MCP-connected agents need enforcement outside the model context?
A: Because the model can describe intent without being the system that enforces it. The harness, server, and tool layer decide whether a request is actually executable, so identity, schema, and provenance checks have to live there. Otherwise the model’s output can still become an authorised external action.
Q: How should teams reduce risk from auto-approval in AI agent workflows?
A: Limit auto-approval to low-risk, tightly bounded calls and remove it from paths that can reach production systems, repositories, or sensitive files. Review whether the agent can reuse approval state across multiple actions, because that is where one decision turns into standing access.
Q: What is the difference between jailbreak resistance and agent authority control?
A: Jailbreak resistance tries to stop the model from producing restricted content or unsafe instructions. Agent authority control stops a valid-looking request from becoming an executable tool call. A system can resist jailbreaks and still fail if its harness authorises the wrong action or exposes the wrong workspace.
Technical breakdown
Where agent authority actually sits in MCP harnesses
An MCP-based agent is not one control plane. It is a chain of model context, tool registration, server identity, and execution policy. If the model can request a tool and the harness can turn that request into a privileged call, then the harness becomes the real security boundary. The article’s examples show why tool name matching, session approval, and parameter handling all matter. A model can be constrained linguistically and still cause an authorised external action if the harness accepts the call. Practical implication: security teams must govern the action path, not just the prompt path.
Practical implication: enforce tool identity, argument validation, and caller provenance outside the model before execution.
Why auto-approval and session approval create standing authority
YOLO mode and session-wide approval turn a one-time decision into persistent execution authority for the life of the session. That is materially different from an isolated permit, because the agent can keep using the same scope without fresh intent checks. In identity terms, the risk is not only excess privilege. It is privilege persistence across a context that was supposed to be ephemeral. Practical implication: session grants need expiry semantics and bounded scope, otherwise the agent inherits standing authority for later actions.
Practical implication: scope approvals to a single call or tightly bounded task window rather than the whole session.
How path traversal and authentication bypass change the trust model
The confirmed bugs in Kimi Code show two classic failure modes: authentication bypass exposed routes without valid bearer-token enforcement, and path traversal let the session filesystem API escape the workspace boundary. Together they collapse the assumption that the harness contains both who may act and what files the action may touch. Once that boundary fails, the agent is no longer operating in a controlled workspace. Practical implication: control must exist at the server and filesystem boundary, not only inside the client workflow.
Practical implication: validate authn and workspace isolation at the server layer, where the action is actually executed.
Threat narrative
Attacker objective: The attacker objective is to turn a model-driven agent session into unauthorised access to tools, routes, and host files through harness weakness rather than prompt manipulation.
- Entry occurred through the agent harness, where a web server bearer token check could be bypassed via percent-encoded API paths and every API route became reachable without authentication.
- Escalation followed when the session filesystem API accepted symbolic links that pointed outside the workspace, letting the session access host files beyond its intended directory boundary.
- Impact was the ability to reach authorised tool paths and local files without the containment the harness was supposed to provide, even without a jailbreak attempt.
Breaches seen in the wild
- Amazon Q Developer extension compromise 2025: An over-scoped CI token let an attacker ship a prompt telling Amazon Q's AI coding agent to wipe files and AWS resources (CVE-2025-8217).
- PocketOS database deletion incident: An AI coding agent found an over-privileged Railway API token in the codebase and deleted PocketOS production data and backups in nine seconds.
Read and download The State of NHI & AI Agent Breach Report 2026, covering 150+ breaches impacting Non-Human Identities including AI Agents.
NHI Mgmt Group analysis
Agent identity fails first at the harness, not at the jailbreak: The article shows that refusal behaviour is only one layer of control. Once the harness can authenticate, approve, and dispatch tool calls, the security problem becomes whether model output can be converted into valid external authority. That shifts the governance question from prompt safety to action authorisation.
Standing approval is the wrong mental model for agent sessions: Session-wide approval and YOLO-style auto-approval treat the agent as if its authority can safely persist across multiple calls. That assumption was designed for short-lived user intent and breaks when the agent reuses the same approval state across changing context. Practitioners should read this as a lifecycle failure of delegated authority, not a UX convenience.
MCP tool permissions need enforcement outside the model context: The article makes clear that tool parameters and intent can sit outside the model’s own permission logic. That means least privilege cannot be inferred from the model’s words or the tool label alone. The practitioner conclusion is that authorisation must inspect the call, not just the conversation.
Agent authority boundary: The useful concept here is the agent authority boundary, the point where model output becomes executable action. The Kimi K3 material shows that this boundary can be crossed by authentication bypass, path traversal, or permissive approval logic before any jailbreak succeeds. Practitioners should treat that boundary as the primary control surface for MCP-connected systems.
Jailbreaks are a distraction when control failures are already documented: Public prompt attacks may draw attention, but the stronger governance signal is the confirmed harness behaviour. If a system can authenticate the wrong call or execute beyond its workspace, then prompt refusal does not contain the blast radius. The field should prioritise action-level controls over model-only safety claims.
From our research library:
- 24,008 unique secrets were exposed in MCP configuration files in 2025 alone, the protocol's first year of widespread adoption, according to the State of Secrets Sprawl 2026.
- Read next: MCP Security Guide
What this signals
Agent authority boundary: The next control question is no longer whether the model can be persuaded, but where a request becomes executable authority. For MCP-connected agents, that boundary needs to live outside the model and inside the server-side enforcement path, because prompt-level refusal does not stop a valid tool call.
Session approvals, YOLO modes, and argument-insensitive permissions all create the same governance problem: one decision can survive long enough to be reused in a different context. That is a lifecycle issue for delegated authority, and lifecycle failures are where agent programmes tend to drift into standing access.
For practitioners
- Separate prompt safety from tool authorisation Map where a model request becomes an executable action and place a control at that boundary. Review MCP server, harness, and client approval logic as separate enforcement points, not one combined trust zone.
- Remove standing approval from agent sessions Treat session approval as temporary task-scoped access, not reusable privilege. Require fresh authorisation for materially different tool calls and expire approvals at the narrowest practical boundary.
- Validate tool arguments outside the model Do not let model-generated intent determine whether a tool call is safe. Check server identity, schema, parameters, and provenance before execution, especially where permission rules ignore arguments.
- Enforce workspace and filesystem containment Test whether symbolic links, path normalization, or filesystem APIs can escape the intended agent workspace. Containment must hold at the server boundary, because harness logic cannot be the only guardrail.
- Audit auto-approval and YOLO modes Identify every path that allows calls to proceed without human review and classify it as standing authority. Disable those paths for repositories, agents, or tools that can reach sensitive systems.
Key takeaways
- The article’s core warning is that agent security can fail in the harness before any jailbreak succeeds, because authenticated tool paths and auto-approval can still produce authorised actions.
- The confirmed Kimi Code bugs show both scale and mechanism, from bearer-token bypass to filesystem traversal outside the workspace.
- The control that matters most is action-level enforcement outside the model, where identity, arguments, and execution context can be checked before a tool runs.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Agentic AI Top 10, OWASP Non-Human Identity Top 10 and MITRE ATT&CK address the attack and risk surface, while NIST AI RMF sets the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP Agentic AI Top 10 | ASI03 — Identity & Privilege Abuse | The article centres on agent authority being reused or misapplied through harness controls. |
| ASI02 — Tool Misuse | The risk is unsafe tool execution through auto-approval and permission gaps. | |
| Recommendation — Apply ASI03 to constrain agent authority at tool-request time, not after the model decides. Map agent tool paths to ASI02 and block calls whose arguments exceed the intended action scope. | ||
| OWASP Non-Human Identity Top 10 | NHI-04 — Insecure Authentication | Confirmed bearer-token bypasses show the agent harness can expose routes without proper authn. |
| NHI-08 — Environment Isolation | Path traversal outside the workspace is an isolation failure at the agent execution boundary. | |
| Recommendation — Harden MCP and agent server authentication so every privileged route is verified before use. Enforce workspace isolation so agent sessions cannot follow links or paths outside their sandbox. | ||
| MITRE ATT&CK | TA0006; TA0008 — Credential Access; Lateral Movement | The article describes credential and action paths that expand access beyond intended boundaries. |
| Recommendation — Hunt for tool-call paths that convert credentials into broader access or movement through connected systems. | ||
| NIST AI RMF | GOVERN — AI Governance and Accountability | The topic is fundamentally about who is accountable for agent actions and approvals. |
| Recommendation — Define governance for who can authorize agent actions and who owns the enforcement boundary. | ||
Key terms
- Agent Authority: The permission an AI agent receives to act on behalf of a verified person. In this model, authority is inherited rather than original, so governance must trace the agent back to the human intent, device context, and current trust state that authorised it.
- Session Approval: A temporary consent state that allows an agent to repeat a class of actions without asking again. In an agentic context, this can behave like standing privilege if the approval window is too broad or the rules do not inspect parameters and context.
- Mcp Permission Matching: A control pattern that decides whether a tool call is allowed based on the tool name or pattern rather than the full argument set. In agentic environments, this can be unsafe because a permitted tool may still perform very different actions depending on its parameters.
- Workspace Containment: The assurance that an agent session can only access files, paths, and resources inside its intended boundary. For agentic systems, containment has to be enforced at the server or filesystem layer, because symbolic links and path handling can otherwise escape the workspace.
What's in the full article
Unosecur's full blog post covers the operational detail this post intentionally leaves for the source:
- Moonshot Kimi Code documentation examples for MCP permissions, hooks, and approval modes
- The specific Kimi Code bugs fixed in version 0.25.0 and how they were documented
- The MCP Gateway control model for identity, JIT access, and per-call enforcement
- Comparisons with other AI client execution paths such as Claude, Cursor, and VS Code
👉 Unosecur's full post covers the Kimi Code bugs, approval modes, and MCP enforcement model
Deepen your knowledge
NHI governance, agentic AI identity, and machine identity lifecycle are core topics in our NHI Foundation Level course, the industry's only accredited NHI security programme. If you are building or maturing an IAM programme, it is worth exploring.
Published by the NHIMG editorial team on August 11, 2026.
Updated on October 7, 2026.
NHI Mgmt Group, the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org