The main break is that authorisation no longer describes the whole action path. Once an agent can generate loops and conditionals, the meaningful risk shifts to what the runtime can do after the initial approval, including repeated calls, fan-out, and state reuse.
When AI agents generate code, what actually changes in the control boundary?
Tool calling keeps the agent inside a bounded set of approved operations. Generated code changes the shape of the workload, because the agent can now express its own control flow, reuse state, branch on intermediate results, and chain actions in ways that are harder to enumerate up front. The control point shifts from “may this tool be invoked?” to “what can this runtime do after the first permitted step?”
That distinction matters in MCP workflows because the initial authorisation event no longer describes the whole execution path. Once code is generated, the meaningful security question is not just whether the agent may call a server, but whether it can expand the blast radius through repeated requests, fan-out, or stateful reuse within a single approved session.
Generated code also weakens simple review assumptions. A human or policy engine can inspect a declared tool call, but it is much harder to pre-audit arbitrary logic that the agent creates on the fly. That makes runtime guardrails, scoped credentials, and action-level policy enforcement more important than a one-time approval check.
Why loops, fan-out, and state reuse are the real breakpoints
Loops and conditionals are the first practical break in the model, because they let an agent turn one approved capability into many executions. A single permitted action can become an unbounded sequence of calls, retries, or branching probes, which changes the risk from discrete access to operational amplification. In practice, that means the relevant unit of control becomes the execution envelope, not the individual request.
State reuse is the second breakpoint. If the generated code can carry forward tokens, identifiers, intermediate outputs, or cached results, then the agent can make later decisions with more context than the original approval assumed. This is where authorisation becomes incomplete: the first decision may be valid, but later steps may inherit trust they did not earn.
That is why MCP Security Guide is useful here, because MCP authorisation is only safe when tokens stay audience-bound and the server does not become a generic execution channel. The same logic applies to MCP authorization for HTTP transports, where the protocol model is built around a resource-server boundary rather than unrestricted token passthrough.
What this means for agent authorisation and runtime policy
Generated code pushes MCP workflows toward per-action policy, not just per-session approval. If an agent can write the logic, then the security model has to evaluate the next action at runtime, including scope, frequency, destination, and whether the action is still consistent with the original intent. This is where least privilege becomes more than a static role design, it becomes an execution constraint.
It also changes what good segregation looks like. A workflow that is acceptable for a single tool invocation may be unsafe once the agent can orchestrate its own sub-steps, because the agent can combine small permissions into a larger effective capability. In that situation, the critical control is to limit what the generated code can do after approval, not only what the initial prompt or plan looked like.
The best fit is often a combination of task-scoped access, short-lived credentials, and hard limits on what runtime code may invoke. NHIMG’s AI Agent Authorisation Guide is directly aligned to that operating model, because it treats authorisation as a per-action decision rather than a broad permission grant. For architecture context, Zero Trust for AI Agents reinforces the same shift: verify the agent, the principal, and the request before each material step.
Risk and Threat Considerations
When agents can generate code, the risk is not only misuse of a tool, but expansion of authorised behaviour into a broader attack surface. Repeated calls, fan-out, and state reuse can turn a bounded approval into excess access, data overreach, or unintended actions against systems the original request did not clearly cover.
Failure mechanism: The agent’s generated logic becomes an execution layer that can loop, branch, preserve state, and reuse credentials or outputs, so the runtime performs more than the initial authorisation captured.
Impact: Defenders lose the ability to reason about the full action path from a single approval, which increases blast radius, makes abuse harder to detect, and raises the chance of privilege amplification or unintended system impact.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Agentic AI Top 10 addresses the attack and risk surface, while NIST SP 800-53 Rev 5, NIST Zero Trust (SP 800-207) and OWASP ASVS set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP Agentic AI Top 10 | ASI03 — Identity & Privilege Abuse | Generated code can expand agent privilege beyond the initial approval. |
| ASI02 — Tool Misuse | MCP workflows can turn approved tools into repeated or unintended actions. | |
| Recommendation — Enforce per-action authorisation to stop agent code from amplifying access. Restrict tool invocation paths and validate each call against intent. | ||
| NIST SP 800-53 Rev 5 | AC-6 — Least Privilege | Code generation can turn narrow access into broader effective capability. |
| Recommendation — Limit credentials and permissions so generated logic cannot exceed need-to-know. | ||
| NIST Zero Trust (SP 800-207) | Zero Trust Architecture | The workflow needs continuous verification after the first approval. |
| Recommendation — Apply continuous verification at each step instead of trusting session start. | ||
| OWASP ASVS | V8 — Authorization | Runtime-generated logic needs authorization checks that survive branching execution. |
| Recommendation — Require authorization decisions for every sensitive action path. | ||
Practitioner Guidance
What to verify: Treat the generated-code step as a separate control boundary. Verify whether the MCP server, downstream API, or sandbox can enforce per-action checks, request limits, and destination restrictions after the first approved call.
What to prioritise: Bound the runtime first, then the prompt. The highest-value controls are short-lived credentials, explicit action scoping, and server-side policy enforcement that can stop repetition or fan-out even when the agent’s code is syntactically valid.
Common mistake: Approving the initial tool call as if it authorises the whole workflow. In generated-code scenarios, the real question is whether the runtime can still constrain later branches, retries, and reused state.
Practitioner takeaway: If the agent can write code, the authorisation decision must move from “may it start?” to “what can it do after it starts?”, because that is where the effective privilege boundary now lives.