TL;DR: Unosecur says Kimi K3 exposed that the real agent security failure sat in the harness, where authentication bypass, path traversal, MCP permission gaps, and auto-approval behaviours could trigger authorised actions without a jailbreak. The boundary problem is not prompt refusal but who can turn model output into external authority.
Editorial analysis by NHI Mgmt Group, based on content published by Unosecur: “The Kimi K3 launch proved the identity problem before anyone tried to jailbreak it”.
Key questions
Q: What breaks when an agent harness treats session approval as standing access?
A: Session approval stops being a narrow control and becomes reusable authority for later tool calls.
Q: Why do MCP-connected agents need enforcement outside the model context?
A: Because the model can describe intent without being the system that enforces it.
Q: How should teams reduce risk from auto-approval in AI agent workflows?
A: Limit auto-approval to low-risk, tightly bounded calls and remove it from paths that can reach production systems, repositories, or sensitive files.
Practitioner guidance
- Separate prompt safety from tool authorisation Map where a model request becomes an executable action and place a control at that boundary.
- Remove standing approval from agent sessions Treat session approval as temporary task-scoped access, not reusable privilege.
- Validate tool arguments outside the model Do not let model-generated intent determine whether a tool call is safe.
Bottom line: The article’s core warning is that agent security can fail in the harness before any jailbreak succeeds, because authenticated tool paths and auto-approval can still produce authorised actions.
What's in the full article
Unosecur's full blog post covers the operational detail this post intentionally leaves for the source:
- Moonshot Kimi Code documentation examples for MCP permissions, hooks, and approval modes
- The specific Kimi Code bugs fixed in version 0.25.0 and how they were documented
- The MCP Gateway control model for identity, JIT access, and per-call enforcement
- Comparisons with other AI client execution paths such as Claude, Cursor, and VS Code
👉 Read Unosecur's analysis of Kimi K3, MCP harness controls, and agent identity risk →
Explore further
View Full Forum → | NHI Foundation Course → | Our Services → | Read the full analysis →
Agent identity fails first at the harness, not at the jailbreak: The article shows that refusal behaviour is only one layer of control. Once the harness can authenticate, approve, and dispatch tool calls, the security problem becomes whether model output can be converted into valid external authority. That shifts the governance question from prompt safety to action authorisation.
A few things that frame the scale:
- 24,008 unique secrets were exposed in MCP configuration files in 2025 alone, the protocol's first year of widespread adoption, according to the State of Secrets Sprawl 2026.
A question worth separating out:
Q: What is the difference between jailbreak resistance and agent authority control?
A: Jailbreak resistance tries to stop the model from producing restricted content or unsafe instructions. Agent authority control stops a valid-looking request from becoming an executable tool call. A system can resist jailbreaks and still fail if its harness authorises the wrong action or exposes the wrong workspace.
👉 Read our full editorial: Kimi K3 exposed the agent identity problem before jailbreaks