Treat approval prompts as necessary but not sufficient. Security teams should require agents to resolve symlinks and show the canonical destination before any write is approved, because the command text alone can hide a config overwrite. The safest control is to make the user approve the real effect, not the literal shell syntax, then limit the agent’s write scope to reduce the blast radius.
Why approval must reflect the real file effect, not just the command text
Approval prompts are useful only if they describe what will actually happen after path resolution. In an AI coding workflow, the same apparent write can land on a different target when a repository contains symlinks, bind mounts, or other path indirection. If the agent approves the literal shell syntax instead of the canonical destination, a harmless-looking edit can become a config overwrite, credential exposure, or destructive change.
The practical standard is simple: resolve the target first, then ask for approval on the resolved object and operation. That shifts review from syntax to consequence, which is the right unit of control when agents can execute file writes autonomously.
What security teams should require before a write is approved
Security teams should require the agent to present the resolved path, the canonical destination, and the expected effect before any write is allowed. That matters most when the repository can redirect the write into a sensitive location outside the apparent working tree. The control should also make it obvious when a write crosses trust boundaries, such as from source code into configuration, secrets, or deployment material.
Where possible, use AI Coding Agents Security Guide as the broader operating model for constraining agent writes, sandboxing execution, and reducing exposure from agent context. For approval design specifically, AI Agent Authorisation Guide is the clearest fit for task-scoped access, per-action decisions, and human approval gates.
A second useful control is to reduce the agent’s write scope so even a mistaken approval cannot reach high-value paths. That can mean limiting writable directories, separating repo workspace from deployment targets, and blocking writes to files whose resolved destination is outside an approved boundary. Approval is then an extra check, not the only barrier.
How to prevent repository tricks from turning approval into a false signal
The core failure mode is path confusion: the reviewer sees one pathname, but the filesystem applies another. Symlinks are the obvious example, but the same risk appears with relative path traversal, generated file targets, and repository-managed indirection that moves a write into a sensitive location. The control fails when the human sees the command and assumes that command text is the effect.
That is why the approval prompt should surface the canonical destination and the resulting operation in plain language. In practice, teams should treat any unresolved or ambiguous target as a deny condition until the agent can show the real destination. When the prompt cannot prove where the write will land, the safest answer is no.
For teams that want a broader security reference on agent behaviour and attack surface, Agentic AI Security Guide provides a useful threat-model lens, and OWASP Agentic AI Top 10 captures the wider class of identity and privilege abuse that can accompany tool-using agents.
Risk and Threat Considerations
When approval is based on the literal command instead of the resolved file target, an attacker or misconfigured repository can steer an agent into overwriting higher-value files than the reviewer intended. The risk is not limited to accidental damage, it also creates a clean abuse path for prompt injection, repository poisoning, and privilege misuse because the filesystem can be used to disguise the real impact of an approved action.
Failure mechanism: The agent approves a write before resolving symlinks or canonical paths, so the visible pathname is different from the effective destination and the write lands in a sensitive location.
Impact: A configuration file, secret-bearing file, or deployment artifact can be overwritten, expanding blast radius and making the approval trail misleading even when the prompt was followed exactly.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Agentic AI Top 10 and OWASP Non-Human Identity Top 10 address the attack and risk surface, while NIST SP 800-53 Rev 5 sets the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP Agentic AI Top 10 | ASI03 — Identity & Privilege Abuse | Approval prompts and resolved write targets control agent privilege abuse. |
| Recommendation — Enforce per-action authorization and deny writes whose resolved destination exceeds the approved scope. | ||
| OWASP Non-Human Identity Top 10 | NHI-05 — Overprivileged NHI | Agent write scope must be limited to prevent oversized file-write blast radius. |
| Recommendation — Constrain agent write privileges to the smallest directory set needed for the task. | ||
| NIST SP 800-53 Rev 5 | AC-6 — Least Privilege | Reducing the agent's writable scope is a least-privilege control for file operations. |
| CM-6 — Configuration Settings | Canonical path checks and workspace boundaries are configuration controls for safe automation. | |
| AU-3 — Content of Audit Records | Audit records should capture the resolved destination, not only the literal command text. | |
| Recommendation — Limit write permissions so approved agent actions cannot reach sensitive paths. Set workspace and filesystem rules so path resolution cannot redirect writes into protected locations. Log the canonical target and approved effect for every agent write operation. | ||
Practitioner Guidance
What to verify: Require the agent to display the canonical path, target ownership, and write scope before approval. If any of those cannot be resolved deterministically, treat the operation as higher risk and do not rely on a prompt that only names the source path.
Decision rule: If the resolved destination can escape the intended workspace, deny the write or force the agent into a narrower sandbox. If the operation stays inside the approved boundary, approval can still be used as a safeguard, but not as the only safeguard.
What good looks like: The reviewer approves the real effect, the agent’s write privileges are small enough to keep a mistake contained, and the path resolution step is visible in logs or audit output.
Practitioner takeaway: Approval should be tied to the filesystem outcome, not the typed command, because path indirection turns a seemingly safe write into a trust and blast-radius problem.
Related resources from NHI Mgmt Group
- How should security teams handle headless AI coding agents that process complete workspaces with repository-local Git configuration?
- How should security teams handle repository files that can run automatically in AI coding tools?
- How should security teams test AI agents after prompts, models, or tools change?
- How do AI agents change the way security teams should handle case lifecycle automation?