TL;DR: Cloudflare’s Code Mode cuts token usage by 32% for a simple task and 81% for a 31-event batch workflow by having agents generate code from MCP server schemas instead of calling tools directly, according to WorkOS. The efficiency gain matters because it shifts MCP design toward hybrid execution models where code generation becomes part of the control surface, not just the model output.
At a glance
What this is: This article shows how Cloudflare's Code Mode changes MCP execution by letting AI agents generate code from server schemas, with reported token savings of 32% on a simple task and 81% on a batch workflow.
Why it matters: It matters because AI agent teams now have to govern not just tool access, but when code generation becomes part of the identity, authorisation, and execution boundary for MCP-backed workflows.
By the numbers:
- Code Mode used 32% fewer tokens for the simple single-event task.
- Code Mode used 81% fewer tokens for the complex 31-event task.
Context
MCP code generation changes the control boundary for AI agent workflows because the agent is no longer only selecting tools, it is also synthesising executable logic from a server schema. That matters for identity governance because the trust decision moves from a single tool call to a generated execution path that can loop, branch, and reuse context.
Cloudflare's demo compares direct MCP tool calling with generated-code execution in a sandboxed Worker. The result is a hybrid pattern in which the agent still reaches the same backend actions, but does so through a more expressive runtime path that reduces repeated calls and context churn.
For AI agent programmes, the governance question is no longer whether the agent can call a tool. It is whether the organisation can safely allow runtime code synthesis against authorised MCP surfaces without losing control over scope, logging, and downstream effect.
Key questions
Q: What breaks when AI agents move from tool calling to generated code in MCP workflows?
A: The main break is that authorisation no longer describes the whole action path. Once an agent can generate loops and conditionals, the meaningful risk shifts to what the runtime can do after the initial approval, including repeated calls, fan-out, and state reuse.
Q: Why do MCP protocol changes create operational risk for AI agent workflows?
A: MCP changes create operational risk because they can alter how agents authenticate, call tools, and use extensions that other systems depend on. If teams assume old behaviour will continue, integrations can fail silently or lose access at runtime. The safer approach is to treat spec updates as dependency changes that require validation, not just documentation review.
Q: How should organisations decide when to use code generation instead of direct MCP tool calls?
A: Use code generation only when the workflow is repetitive, schema-driven, and stable enough that a looped runtime path will reduce friction without expanding scope. For one-off actions or loosely defined tasks, direct tool calls remain easier to govern and audit.
Q: What should security teams look for when reviewing sandboxed agent execution?
A: They should check the permissions inherited by the sandbox, the backend actions it can reach, and whether execution logs separate generated logic from direct model output. If those elements are blurred, the organisation will not know what the agent actually did.
Technical breakdown
Direct tool calling versus generated code in MCP
In standard MCP usage, the agent selects a tool, sends a request, receives a response, and repeats that cycle for each step. Code Mode changes the execution model by turning the MCP server schema into source material for generated code, which can then perform loops, conditional logic, and repeated API calls inside one runtime. The practical effect is fewer round trips and less context overhead, especially for repetitive workflows. The architectural trade-off is that the security boundary is no longer just the tool endpoint. It now includes the generated program, the sandbox, and the permissions that code inherits at execution time.
Practical implication: Model generated code as part of the authorised attack surface and review its runtime permissions separately from the MCP server itself.
Why batch workflows amplify token savings in MCP
Batch tasks expose the weakness of pure tool calling because every event, record, or object may require a separate call and response. That creates token churn, repeated state tracking, and more opportunities for the model to drift or lose context. Generated code can compress that sequence into a loop that handles many similar actions with one reasoning step and many execution steps. For agent teams, the important point is not only efficiency. It is that expressive execution can change where errors surface, making correctness depend on schema quality, code generation quality, and sandbox behaviour rather than on a single prompt-response exchange.
Practical implication: Use batch-heavy workflows as the first candidate for code-generation designs, but validate schema fidelity and sandbox behaviour before expanding scope.
Sandboxed execution becomes part of the governance model
Cloudflare's architecture runs generated code inside a sandboxed Worker, which keeps the execution isolated while still allowing access to the MCP server. That separation is critical because it makes the runtime a governed control point rather than an informal extension of the model. In identity terms, the MCP server is no longer the only object being authorised. The generated worker identity, its execution lifetime, and its permitted egress paths all become relevant. For NHI governance, this looks less like a prompt optimisation and more like a new non-human execution tier that needs explicit policy boundaries.
Practical implication: Define what the sandboxed worker may access, persist, and call before treating code generation as safe for production workloads.
Threat narrative
Attacker objective: The objective is to execute high-volume, schema-driven actions efficiently while preserving authorised access to backend services.
- Entry begins when an agent receives access to an MCP server definition and generates executable code from that schema instead of making direct tool calls.
- Escalation occurs when the generated code gains loop logic, conditional branching, and repeated invocation capability inside the sandboxed runtime.
- Impact emerges when that runtime can carry out large batches of authorised actions with lower token cost and fewer interactive checkpoints than direct tool calling.
Breaches seen in the wild
- Anthropic GTG-1002 AI espionage campaign: A state-sponsored group ran Claude Code agents to attack about 30 organisations, harvesting and reusing credentials at machine speed.
Read and download The State of NHI & AI Agent Breach Report 2026, covering 150+ breaches impacting Non-Human Identities including AI Agents.
NHI Mgmt Group analysis
Code generation is now part of the MCP control surface: When an AI agent turns a server schema into executable code, the governance problem changes from single-call authorisation to runtime execution authorisation. That matters because loops, conditionals, and repeated API calls can amplify the effect of one approved interaction. Practitioners should treat generated code as governed non-human execution, not as a harmless optimisation.
Ephemeral token savings do not remove authorisation risk: The efficiency gain is real, but it does not answer who approved the downstream actions, how scope is bounded, or what prevents generated logic from crossing intent boundaries. In practice, the more expressive the runtime, the more governance has to move upstream to schema quality, sandbox policy, and execution logging. The implication is that agent performance tuning and identity governance now share the same control plane.
Runtime code synthesis creates an identity blast radius problem: A task that used to require 31 visible tool calls can now be compressed into one generated program that fans out internally. That makes blast radius harder to see with traditional access review or prompt review techniques because the interesting decision happens after authorisation, inside execution. The practitioner conclusion is that access scope alone is no longer the full story; execution shape now matters.
Hybrid MCP patterns will dominate because they separate simple calls from complex orchestration: Direct tool invocation remains appropriate for simple actions, while generated code is better suited to repetitive multi-step workflows. The governance challenge is that both paths can hit the same backend, but their risk profiles are not identical. Teams should assume their MCP architecture will become mixed by default and build controls that distinguish between direct action and generated orchestration.
Sandboxing is necessary, but not sufficient, for agentic code execution: A sandboxed Worker constrains blast radius, yet it does not automatically validate whether the generated logic is appropriate, least-privileged, or aligned with the original request. The field should treat sandboxing as the last containment layer, not the primary trust decision. Practitioners need policy around what kinds of MCP tasks may be promoted from tool calls to generated code.
From our research library:
- 24,008 unique secrets were exposed in MCP configuration files in 2025 alone, the protocol's first year of widespread adoption, according to the State of Secrets Sprawl 2026.
- Read next: MCP Security Guide
What this signals
Code-generated execution broadens the governance boundary: The critical question is no longer only who can call an MCP tool, but which generated runtimes are allowed to interpret schemas and fan out into backend actions. That pushes policy toward execution-time control, especially where a single request can become a multi-step program.
Agentic code paths will force new review models: Traditional access review assumes a stable entitlement can be certified after the fact. When the agent synthesises a program on demand, the more useful control is pre-authorisation of task classes and sandbox conditions rather than retrospective review of every internal step.
Hybrid MCP design is becoming the practical default: Simple tool calls and generated orchestration solve different problems, so mature programmes will need separate governance for each. The challenge is to preserve auditability and least privilege while still allowing the runtime to choose the cheaper path when complexity demands it.
For practitioners
- Define code-generation eligibility for MCP tasks Classify which workflows are simple enough for direct tool calls and which are repetitive enough to justify generated code. Reserve code execution for tasks where the batch structure materially reduces repeated calls and where the schema is stable enough to be trusted.
- Treat sandboxed workers as governed execution identities Assign explicit policy to the worker runtime, including allowed MCP targets, egress limits, and logging requirements. Do not assume isolation alone makes the generated program safe.
- Review MCP schemas as security-relevant inputs Check whether schemas expose actions that are broader than the business task actually requires. A schema that is too permissive will let generated code do more than direct tool calling would have made obvious.
- Instrument code-generation paths separately from tool calls Track generated programs, their execution IDs, and their backend effects so that reviews can distinguish direct action from orchestrated action. Without that separation, auditability collapses into generic model telemetry.
Key takeaways
- Code generation changes MCP from a simple tool-calling pattern into a governed runtime path with broader execution implications.
- The efficiency gains are strongest in batch workflows, where repeated tool calls otherwise consume context and tokens.
- Security teams should classify which MCP tasks may use generated code and define the sandbox, logging, and scope rules before production use.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Agentic AI Top 10 and OWASP Non-Human Identity Top 10 address the attack and risk surface, while NIST CSF 2.0 sets the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP Agentic AI Top 10 | ASI02 — Tool Misuse | The article centres on agents choosing a different execution path for MCP actions. |
| ASI03 — Identity & Privilege Abuse | Generated code inherits runtime authority that must be governed as privilege. | |
| Recommendation — Constrain agent tool-use pathways so generated code cannot exceed approved action scope. Treat runtime-authorised code paths as privilege-bearing and review their execution boundaries. | ||
| OWASP Non-Human Identity Top 10 | NHI-04 — Insecure Authentication | MCP code execution still depends on trusted identity and authorisation boundaries. |
| NHI-05 — Overprivileged NHI | Sandboxed workers and MCP targets can become overpowered if scope is not narrowed. | |
| Recommendation — Bind MCP execution to explicit authentication and authorisation boundaries before enabling generated code. Reduce worker and MCP scope to the minimum required for each task class. | ||
| NIST CSF 2.0 | PR.AA-05 — Access Permissions, Entitlements and Authorizations | The article is fundamentally about who or what is permitted to execute actions through MCP. |
| Recommendation — Define and enforce permissions for both direct tool calls and generated execution paths. | ||
Key terms
- Generated code execution: A runtime model where the agent writes code from a tool schema and the environment executes it inside a sandbox. This compresses repeated actions into one execution block, which improves efficiency but requires stronger governance around what the code can do once it starts running.
- MCP tool calling: MCP tool calling is the act of an AI agent invoking an external capability through the Model Context Protocol. It lets the agent request data, actions, or functions from connected tools in a structured way. The protocol standardizes how requests, permissions, and responses are exchanged between the agent and the tool.
- Sandboxed Worker: An isolated execution environment used to run generated code with bounded permissions and controlled network access. For identity teams, the sandbox is the place where agent authority becomes operational, so it must be treated as a governed runtime, not just an implementation detail.
- Execution boundary: The point at which an authorised task turns into a real system change, such as writing data, deleting records, spending money, or invoking a downstream tool. In AI governance, controlling the execution boundary matters more than simply approving access, because harm occurs when actions are allowed to complete unchecked.
Deepen your knowledge
NHI governance, agentic AI identity, and machine identity lifecycle are core topics in our NHI Foundation Level course, the industry's only accredited NHI security programme. If you are building or maturing an IAM programme, it is worth exploring.
Published by the NHIMG editorial team on June 7, 2026.
Updated on October 7, 2026.
NHI Mgmt Group, the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org