Join our Newsletter — 33% off our NHI Course

What breaks when AGENTS.md is treated as trusted input for coding agents?

The control boundary breaks because the agent can execute repository instructions before the user task and without a separate human approval step. That turns a file meant for project guidance into an instruction channel with the same practical effect as code execution. Security teams should treat that file as untrusted until reviewed and policy-cleared.

What AGENTS.md Actually Changes in a Coding Agent Workflow

AGENTS.md is not just project documentation when a coding agent is allowed to ingest it as instruction. It becomes part of the agent’s control plane: the file can shape task execution, tool use, and repository behavior before the user’s request is fully mediated. That matters because the trust model has shifted from “read for context” to “execute as guidance,” which is a different and much sharper boundary.

That shift is especially visible in ai coding assistant and terminal agents, where repository-local instructions can influence how the agent interprets the work, what files it touches, and which commands it runs. NHIMG’s AI Coding Agents Security Guide covers the broader class of risks that appear when developer tools treat repository content as actionable input rather than passive text.

Once that boundary is crossed, the file is no longer just a convenience for maintainers. It can become a durable instruction channel that survives review gaps, copy-paste workflows, and automation defaults, which is why project-local guidance deserves the same skepticism as any other externally supplied instruction source until it has been policy-approved.

Which Security Assumptions Break First

The first assumption that breaks is human mediation. If the agent can consume AGENTS.md before the user task and without a separate approval step, the repository can steer behavior before a person has confirmed the request is safe. The second assumption is instruction provenance: developers may expect only the user to author the task, but the repo can now contribute commands, constraints, and priorities.

This creates practical overlap with prompt injection and indirect instruction abuse. A poisoned repository file can alter execution paths even when the user prompt itself is benign. NHIMG’s Gemini CLI prompt injection flaw 2025 is a good example of how hidden instructions in project content can push a coding agent into silent command execution and secret exposure.

It also raises the stakes for authorization. If the agent is already connected to cloud tokens, source control credentials, or local developer privileges, the file can influence actions that have real blast radius. NHIMG’s AI Agent Authorisation Guide is relevant here because the core control problem is not just what the agent reads, but what it is allowed to do after reading it.

How Teams Should Treat AGENTS.md in Practice

Security teams should classify AGENTS.md as untrusted repository input until the content has been reviewed, versioned, and explicitly allowed by policy. The safe default is to separate descriptive repository guidance from executable agent instructions, then decide which paths are permitted to influence tool use, shell access, or file modification.

What to verify: confirm whether the agent treats AGENTS.md as pre-task instruction, post-task guidance, or a hard policy source. If it can affect execution before human approval, require review gates and scope limits before enabling it in production workflows.

Decision rule: if a repository file can change commands, tool calls, or file writes, treat it as an untrusted input surface and apply the same review discipline you would use for other instruction-bearing artifacts. If it only provides static documentation, its security impact is lower, but it still needs provenance controls in shared repositories.

Common mistake: assuming “repo-local” means “safe.” Local origin does not imply trusted intent, especially in pull-request driven environments, fork workflows, or agentic IDEs that run with persistent credentials.

Risk and Threat Considerations

When repository instructions are treated as trusted, the main risk is control-plane confusion: a file intended for guidance can steer an agent with the practical effect of code execution. That widens the attack surface to include poisoned repos, compromised contributors, and malicious changes that are hard to spot in review.

Failure mechanism: the agent ingests project instructions before user intent is fully constrained, then follows those instructions with access to tools, credentials, or file-write capabilities. That can lead to unauthorized command execution, secret exposure, or destructive repository actions.

Impact: the outcome is not just bad documentation, but delegated execution with insufficient separation between human approval and machine action. At scale, that can turn routine developer tooling into a repeatable path for supply-chain abuse or insider-style manipulation.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

OWASP Agentic AI Top 10 and MITRE ATT&CK address the attack and risk surface, while NIST SP 800-53 Rev 5 sets the governance and control requirements practitioners need to meet.

Framework Control / Reference Relevance
OWASP Agentic AI Top 10 ASI03 — Identity & Privilege Abuse AGENTS.md can steer agent action and privilege use before user approval.
ASI02 — Tool Misuse Trusted repo instructions can redirect tools, shell commands, and file writes.
ASI01 — Agent Goal Hijack A malicious AGENTS.md can redirect the agent away from the user’s intended task.
Recommendation — Require per-action approval when repository instructions can alter agent authority. Constrain tool execution so repository text cannot trigger unsafe actions. Validate task intent before allowing repository instructions to shape execution.
MITRE ATT&CK T1204 — User Execution The attack relies on manipulating an operator or agent into running attacker-supplied instructions.
Recommendation — Hunt for repository-supplied instruction paths that lead to execution.
NIST SP 800-53 Rev 5 AC-6 — Least Privilege Limiting agent permissions reduces harm if repository instructions are abused.
Recommendation — Scope agent permissions to the minimum needed for the task.

Practitioner Guidance

What to prioritise: put approval boundaries around any agent that can read repository instructions and then act on them. The highest-value control is limiting what the agent may do before a human has validated the task scope and the instruction source.

What good looks like: the agent can read AGENTS.md for context, but cannot treat it as authoritative for commands, secrets handling, or writes unless the file has been explicitly reviewed and allowed. A safe workflow also makes instruction provenance visible in logs so reviewers can distinguish user intent from repository influence.

Practitioner takeaway: the key decision is whether the file informs the agent or governs it. Once it governs execution, you have moved from documentation hygiene into authorization and control-boundary design.