The control becomes documentary rather than operational. Teams may believe security guidance is enforced because it is written down, but the agent is still free to weight other cues, use stale instructions, or miss the rule entirely when the file is large and the task is specific.
When AGENTS.md Becomes Policy Theater Instead of Control
AGENTS.md can help when it is a shared, current source of intent, but it breaks down as a control the moment teams assume the file itself enforces behavior. The gap is between documented guidance and runtime governance: an agent can still follow stronger or newer cues, ignore buried instructions, or act inconsistently when the instruction set is long and the task is narrow.
That is why instruction files should be treated as one input to agent behavior, not as proof of constrained behavior. If the control depends on an agent reading, remembering, and prioritising the right text every time, then the control is already softer than most teams assume. For coding agents, the relevant risk is not only compliance drift, but also secret exposure, over-scoped actions, and instruction conflicts that surface only in specific task contexts.
In practice, the control fails most obviously when the file is large, the task prompt is specific, or multiple instruction sources compete. The agent may comply with the visible request while skipping the security rule, or it may apply an outdated instruction because the file was not refreshed or was not the highest-priority source in the runtime context.
Why the Failure Is Operational, Not Just Documentary
Security teams often confuse written policy with enforced policy. With agent instruction files, that confusion matters because the model does not “inherit” control authority from the filename, it only processes text according to context, retrieval, and prompt hierarchy. If the security requirement is not expressed in a way the agent reliably sees and retains, the rule can be effectively absent during execution.
That distinction matters most in workflows where the agent can touch code, secrets, package choices, or build outputs. A file that says “do not expose credentials” is not the same as a system that prevents the agent from seeing them, limits its tool scope, or blocks unsafe actions. For deeper agent security patterns, teams should anchor their understanding in the AI Coding Agents Security Guide, which covers secrets in context, sandboxing, and over-scoped tokens.
When AGENTS.md is treated as a control, ownership also becomes fuzzy. The file often lives with developers, while the real control needs engineering, platform, and security agreement on precedence, scope, and verification. Without that, teams end up with a rule that is easy to review and hard to trust.
For agent-specific governance, the more relevant question is whether the system can bound the agent’s authority independent of instruction quality. The AI Agent Authorisation Guide is useful here because it separates task-scoped approval from loose textual instruction and forces per-action decisions.
What Good Looks Like When AGENTS.md Is Only One Layer
The strongest pattern is to make AGENTS.md descriptive, then back it with controls that do not depend on perfect prompt compliance. That usually means narrowing tool access, bounding credentials, and deciding which actions require explicit approval even if the file says the agent should behave safely.
Teams should also verify whether the agent actually consumes the file at the point of action, not just at the start of a session. If the agent uses long contexts, multiple repositories, or generated sub-tasks, the security rule may be present but no longer influential by the time the risky action occurs.
A practical benchmark is whether the same safety rule survives file length, task specificity, and prompt variation. If a rule only works when it is short, recent, and central to the prompt, then it is guidance, not a dependable control. For implementation patterns that combine identity, access, and runtime governance, the Agentic AI Security Policy Template is a useful starting point because it covers registration, access, human oversight, monitoring, and retirement.
That is also where monitoring becomes essential. If the agent can act, then the team needs evidence of what it saw, what it selected, and what it did. Without logs or audit trails, you cannot tell whether the file failed, the agent ignored it, or the task was simply outside the control boundary.
Risk and Threat Considerations
When teams overtrust AGENTS.md, they create a false sense of policy enforcement. The risk is that unsafe behavior looks governed on paper while remaining fully available at runtime, especially when instruction priority is ambiguous or the agent can be steered by competing context.
Failure mechanism: The agent treats the file as advisory context, not as an enforced control, so instruction conflicts, stale content, or context truncation can let unsafe actions proceed.
Impact: Security guidance can be bypassed without a visible control failure, leading to secret exposure, excessive tool use, unsafe code changes, or inconsistent behavior across similar tasks.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Agentic AI Top 10 addresses the attack and risk surface, while NIST SP 800-53 Rev 5 and NIST Zero Trust (SP 800-207) set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP Agentic AI Top 10 | ASI03 — Identity & Privilege Abuse | AGENTS.md can fail when agent authority exceeds intended instructions. |
| ASI02 — Tool Misuse | The issue concerns agents taking unsafe tool actions despite written guidance. | |
| Recommendation — Bind each risky agent action to explicit authorization and least privilege. Restrict tool access and require approval for high-impact operations. | ||
| NIST SP 800-53 Rev 5 | AC-6 — Least Privilege | Written instructions do not enforce bounded authority at runtime. |
| AU-2 — Event Logging | Trusting AGENTS.md needs evidence of what the agent actually did. | |
| Recommendation — Limit agent permissions to the minimum needed for the task. Log agent decisions and actions so instruction failures are observable. | ||
| NIST Zero Trust (SP 800-207) | Zero Trust Architecture | The control problem is trusting a file instead of verifying runtime behavior. |
| Recommendation — Verify each agent action at runtime instead of assuming repo instructions are enforced. | ||
Practitioner Guidance
What to verify: Confirm whether AGENTS.md is actually loaded, weighted, and retained at the moment a risky action is chosen, not just whether the file exists in the repo.
Common mistake: Treating repository instructions as if they are policy enforcement. If the agent can still reach the action with broader permissions, the file is advisory and must be paired with real runtime controls.
Decision rule: If the instruction must never be missed, implement a control that fails closed outside the file, such as scoped permissions, approvals, or tool restrictions, and use AGENTS.md only as supporting guidance.
Practitioner takeaway: The file can shape behavior, but only runtime constraints prove control, so trust AGENTS.md for guidance, not for security assurance.
Related resources from NHI Mgmt Group
- How should security teams manage permissions for AI agents?
- How should security teams govern AI agents that use OAuth access?
- How should security teams limit the risk from AI agents that have access to production systems?
- How should security teams govern AI agents that can access enterprise systems?
Deepen Your Knowledge
Free weekly newsletter
Subscribe to the NHI & AI Identity Journal
The latest on NHI and Agentic AI security – articles, research, breaches, news and events every week.
Bonus 33% off our NHI Course when you subscribe.
Reviewed and updated by the NHIMG editorial team on October 11, 2026.
NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org