Treat the sandbox as an enforced system boundary, not a prompt instruction. Validate the resolved filesystem path after all variable expansion and command parsing, then block any command that touches data outside the allowed directory. Static string checks are brittle because attackers can hide traversal sequences in variables or other runtime transformations that only appear after validation.
Why path validation beats string checks in shell-facing agent sandboxes
When a natural-language system turns intent into shell commands, the real control is whether the resolved command can reach anything outside the intended working area. That means the sandbox must be enforced after expansion, parsing, and path resolution, not guessed from the raw text. A brittle precheck can miss traversal that only appears once variables, quoting, or substitutions are resolved.
For agentic systems, this is a boundary problem as much as an input problem. The agent may be operating with delegated authority, but the shell runner still needs a hard policy on what paths are reachable and what file operations are allowed. If the boundary is only implied by prompts or instructions, a malformed command can turn an ordinary workflow into an escape.
Path validation also needs to be outcome-based. The question is not whether the command text looks safe, but whether the final filesystem target is inside the allowed directory after every transformation the shell will apply. That is why path canonicalisation, symlink awareness, and post-parse enforcement matter more than regex-style deny lists.
Where escapes usually happen
The common failure mode is validating the wrong string. Attackers can place traversal markers in a variable, hide path components behind command substitution, or rely on shell expansion to change the meaning of a command after a superficial check has passed. If the system approves the pre-expansion string, it may later execute a command that points well outside the sandbox.
Another weak point is assuming that a directory prefix alone proves containment. A path that begins inside the sandbox can still resolve elsewhere through symlinks, relative components, mount boundaries, or renamed working directories. The control must compare the fully resolved target against the permitted root, not just the visible text of the request.
Agentic workflows make these mistakes easier to weaponise because the command is often assembled from many steps: model output, planner output, tool output, environment variables, and local runtime state. Security teams should treat each step as a place where the final command can diverge from the apparent one. The right question is whether the command remains safe after all runtime transformations, not whether it was safe in a template.
How to design a usable containment model
The most reliable pattern is to make the sandbox enforce a small set of file operations and a narrow root, then reject anything that resolves outside that root. That should happen at execution time, close to the syscall or process boundary, so the policy sees the actual path the shell will use. If a command cannot be evaluated against the final resolved path, it should fail closed.
Security teams should also decide whether the agent needs general shell access at all. If the workflow only needs a few known file operations, a purpose-built tool is safer than broad command execution because it shrinks the parsing surface and reduces the number of expansion rules that can be abused. The less shell syntax the agent can influence, the less room there is for hidden path manipulation.
For larger agentic systems, the containment model should be paired with logging and review of blocked attempts. That gives teams evidence that the boundary is actually being exercised and helps distinguish a noisy misconfiguration from a real escape attempt. AI Coding Agents Security Guide covers why sandboxing must be treated as part of the agent execution model, not an optional hardening step. Agentic AI Security Guide is useful for placing file-access controls alongside the wider agent threat model. Zero Trust for AI Agents reinforces the same design principle: verify the request against policy at the point of action.
Risk and Threat Considerations
Sandbox escapes are dangerous because they convert a bounded automation task into arbitrary host access. In agentic systems, that can expose source code, secrets, configuration files, or sibling workloads even when the original prompt looked harmless. The threat is strongest when the system trusts pre-expansion text or assumes the sandbox is only advisory.
Failure mechanism: An attacker places traversal or path-rewrite logic into variables, substitutions, or symlinked paths so the command resolves outside the intended directory after validation, then the shell executes the escaped path with the agent's authority.
Impact: File disclosure, unauthorized modification, secret theft, or follow-on execution outside the sandbox can occur, and the same weakness can become a stepping stone to broader compromise if the agent runs with access to sensitive tooling or credentials.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Agentic AI Top 10 addresses the attack surface, NIST SP 800-53 Rev 5 and CIS Controls v8 set the technical controls, and ISO/IEC 27001:2022 defines the regulatory obligations.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP Agentic AI Top 10 | ASI05 — Unexpected Code Execution | Shell escape paths in agents are code-execution boundary failures. |
| Recommendation — Enforce runtime policy checks before any agent-generated command reaches execution. | ||
| NIST SP 800-53 Rev 5 | AC-3 — Access Enforcement | Sandbox containment is an access decision over filesystem reachability. |
| SI-10 — Information Input Validation | Runtime path handling requires validation after expansion and parsing. | |
| Recommendation — Enforce filesystem and command boundaries at the execution layer. Validate resolved command inputs after all transformations. | ||
| CIS Controls v8 | CIS-4 — Secure Configuration of Enterprise Assets and Software | Hardened execution environments reduce escape opportunities from unsafe defaults. |
| Recommendation — Harden the sandbox and restrict the runtime's reachable filesystem. | ||
| ISO/IEC 27001:2022 | A.8.9 — Configuration management | Sandbox boundaries depend on controlled runtime configuration and enforced settings. |
| Recommendation — Lock down sandbox configuration and review changes to execution boundaries. | ||
Practitioner Guidance
What to verify: Test the control against the final resolved path, not the string you expected the model to generate. Use cases should include variable expansion, nested quoting, symlinks, relative traversal, and any runtime rewriting the shell performs.
Common mistake: Teams often block obvious ../ sequences and stop there. That is not enough if a later transformation can reintroduce the escape path after the check has already passed.
What good looks like: A command that resolves outside the allowed root is rejected consistently, regardless of how it was constructed, and the agent cannot rely on prompt wording to override that decision.
Practitioner takeaway: Treat the sandbox as a runtime enforcement boundary with path resolution at the point of execution; if policy is decided before expansion, the control is not actually controlling the command.
Related resources from NHI Mgmt Group
- How should security teams prevent communication poisoning in agentic AI systems?
- How should security and AI teams design agentic systems so smaller language models handle routine work without weakening reliability?
- How should security teams contain prompt injection in agentic systems?
- What do security teams get wrong about least privilege for agentic systems?
Deepen Your Knowledge
Free weekly newsletter
Subscribe to the NHI & AI Identity Journal
The latest on NHI and Agentic AI security – articles, research, breaches, news and events every week.
Bonus 33% off our NHI Course when you subscribe.
Reviewed and updated by the NHIMG editorial team on September 30, 2026.
NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org