Sandboxing code execution means the isolated environment starts only when the agent needs to run code, so text-only turns stay cheap and high concurrency remains possible. Wrapping every session in a container isolates everything, but it also adds cost and reduces capacity. For most production agents, the right control is on-demand isolation tied to execution, not to conversation length.
Why On-Demand Sandboxing Differs from Full-Session Container Isolation
Sandboxing code execution isolates the risky part of an agent’s work only when code is actually about to run. That keeps ordinary text turns lightweight and preserves concurrency. Wrapping every session in a container isolates the whole interaction, which is stronger from a perimeter standpoint, but it also consumes more compute and can reduce throughput in production systems.
That distinction matters because “more isolation” is not automatically “better control.” The right question is whether the risk lives in the execution step, the whole session, or both. If the agent spends most of its time reasoning, retrieving, or drafting, full-session isolation can be an expensive way to protect a narrow execution moment.
Sandboxing is therefore a control pattern, not just an infrastructure choice. It narrows the blast radius of code execution by creating an isolated runtime only for the dangerous action, then tearing it down when the action ends. That makes it easier to scale agents that handle many non-executing turns, while still keeping a hard boundary around code that could touch files, networks, or local system state.
What Changes When the Boundary Covers the Whole Session
Full-session containerization changes the unit of isolation from “the execution event” to “the conversation.” Everything the agent does, including text-only turns, happens inside the container boundary. That can simplify some trust decisions, but it also means the platform must provision and maintain a heavier environment for every session, even when no code ever runs.
The practical trade-off is capacity. If you isolate every session, you usually pay for startup overhead, memory footprint, storage, and orchestration at a much larger scale than an on-demand sandbox. For high-volume agents, that can turn a sensible defense into a bottleneck. The better design is often to isolate the part that needs containment, not the entire dialogue path.
There is also a difference in failure scope. A session container can help contain state, temp files, and local artifacts across multiple turns, but it does not automatically solve unsafe code behavior if the container is still allowed broad network or filesystem access. Conversely, a code sandbox can be very effective if the main concern is untrusted execution, even if the broader session remains outside the container.
How to Choose the Right Control Boundary
The selection rule is straightforward: isolate what can cause material harm, and keep the cheaper path for everything else. If the dangerous moment is code execution, use on-demand sandboxing tied to that event. If the agent must preserve a long-lived working state with meaningful local side effects across many turns, session-level containerization may be justified. The NIST SP 800-190 Container Security guidance is useful here because it treats container image, runtime, and orchestration boundaries as distinct risk surfaces rather than one generic control.
For production teams, the key design question is blast radius versus operational cost. On-demand sandboxing usually wins when most sessions are non-executing, because it preserves throughput and reduces idle overhead. Full-session containers make more sense when the session itself is the trust boundary, such as when the agent accumulates local artifacts, stateful tool outputs, or sensitive intermediate files that must remain contained between turns.
A useful way to test the design is to ask what happens when the agent never executes code. If the answer is “we still paid for isolation anyway,” the control may be too coarse. If the answer is “we would have lost state or exposed local artifacts without the container,” then session-level isolation is doing real work. The best architecture matches isolation cost to actual risk, not to the length of the conversation.
Risk and Threat Considerations
The main risk is over-isolating by default or under-isolating at the moment code actually runs. Too little isolation can let untrusted code touch files, network endpoints, or secrets outside the intended boundary. Too much isolation can create capacity pressure, which encourages teams to weaken controls, skip isolation for performance, or reuse environments unsafely.
Failure mechanism: The control fails when the isolation boundary does not match the event that creates risk. If code execution is the only dangerous step, but the platform containers every turn, the architecture wastes resources without materially improving safety. If code can execute outside a real sandbox, the agent can still reach data or systems that were supposed to be contained.
Impact: Misaligned isolation either increases exposure or reduces service quality. In the first case, you get a path for untrusted code to act with too much reach. In the second, you get higher cost, lower concurrency, and pressure to relax the very safeguards the design was supposed to enforce.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
NIST SP 800-53 Rev 5 and CIS Controls v8 set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| NIST SP 800-53 Rev 5 | SC-7 — Boundary Protection | Defines containment boundaries for untrusted agent execution and session isolation. |
| CM-7 — Least Functionality | Supports disabling unnecessary capabilities in sandboxes and session containers. | |
| AC-6 — Least Privilege | Limits what sandboxed code or full-session containers can access or do. | |
| Recommendation — Enforce boundary controls to confine untrusted execution to the smallest viable scope. Remove unneeded services and permissions from the isolated runtime. Grant the isolated environment only the minimum permissions required. | ||
| CIS Controls v8 | CIS-4 — Secure Configuration of Enterprise Assets and Software | Container and sandbox hardening depends on secure runtime configuration. |
| CIS-8 — Audit Log Management | Session and sandbox events need logging to understand execution and containment. | |
| Recommendation — Harden container and sandbox defaults before allowing execution. Log sandbox launches, code execution, and containment failures. | ||
Practitioner Guidance
What to prioritise: Draw the boundary around the action that needs containment, not around the entire conversation. If only code execution is risky, build a fast sandbox that can be launched and destroyed per execution event.
What to verify: Confirm that the sandbox actually limits the capabilities you care about, especially filesystem access, outbound network reach, and inherited credentials. A container is not automatically a safe boundary if it still has broad permissions.
Trade-off: Treat full-session containers as a state-management choice as much as a security choice. They are justified when the session itself carries meaningful local state, but they should not become the default just because they sound stricter.
Practitioner takeaway: In most agent systems, the best control is the narrowest one that still contains the dangerous action, because that preserves both security and scale.
Related resources from NHI Mgmt Group
- What is the difference between code scanning and runtime identity monitoring?
- What is the difference between eBPF and proprietary sandboxing for secure code execution in operating systems?
- What is the difference between sanitising LLM output and sandboxing code execution?
- What is the difference between execution sandboxing and agent authorization?
Deepen Your Knowledge
Reviewed and updated by the NHIMG editorial team on September 30, 2026.
NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org