They create risk because the assistant may treat those instructions as inherited authorisation, not as untrusted text. That can turn a harmless-looking repo change into a durable mechanism for command generation, credential theft, or other actions that would otherwise be refused.
Why project-scoped instructions become a control boundary problem
Project-scoped instruction files look like configuration, but agentic coding tools often treat them as part of the active execution context. That matters because the tool may blend repository text with user intent and system policy, then generate actions from the combined prompt. Once untrusted text can influence tool use, the repository itself starts shaping what the assistant is willing to do.
The risk is not the instruction file alone, but the fact that agentic tools can follow it across many interactions and tasks. A malicious or simply overbroad instruction can persist, steer future completions, and create a durable path for unsafe behavior even when the user never explicitly asked for it.
That is why project-scoped instructions are less like comments and more like an untrusted policy input. In practice, they sit close to the boundary between developer convenience and delegated action, which is exactly where abuse becomes possible.
How inherited authority turns a harmless change into a risky one
Agentic coding tools are especially exposed when they infer authority from location rather than provenance. If a file in the repository is treated as trusted guidance, then any collaborator, dependency, or compromised branch that can edit that file may be able to reshape the tool’s behavior. The issue is not just prompt injection in the abstract, but the false assumption that project-local text deserves inherited authority.
Once that assumption exists, the instructions can become a mechanism for escalation by social engineering the tool rather than the human reviewer. The model may be nudged to reveal secrets, expand scope, call tools it should not call, or make changes that look routine but actually create a covert control path for later abuse.
For coding tools that can read files, run commands, or modify code, the boundary error is severe. A project instruction that seems to improve productivity can quietly become a standing source of influence over execution, review, and commit behavior.
What this changes for secure use of agentic coding tools
Project instructions should be treated as data that is parsed, not authority that is inherited. The safest operating model is to assume repository text can be attacker-controlled unless it is explicitly vetted, especially in shared repos, fork-based workflows, or projects that ingest external contributions.
Teams should distinguish between harmless convenience guidance and text that can alter tool actions. Instructions that affect shell use, dependency installation, credential handling, file write scope, network access, or code generation deserve the same caution as any other untrusted input that can influence privileged behavior.
Reviewers should also ask whether the tool has a clear place to express policy outside the repo, such as centrally managed defaults or action approval gates. When local instructions can override safer defaults, the assistant’s behavior becomes harder to reason about and harder to audit.
Risk and Threat Considerations
Project-scoped instructions are risky because they can be used as a persistence and influence channel inside the assistant’s working context. If an attacker can alter them, the tool may repeatedly follow hostile guidance without any further user interaction.
Failure mechanism: The assistant mistakes repository text for trusted operational policy, then uses that text to justify actions, disclosure, or tool calls that should have required separate authorization.
Impact: The result can be secret exposure, unauthorized code changes, unsafe command execution, or a durable prompt-injection style control path that survives normal task boundaries.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Agentic AI Top 10 and MITRE ATT&CK address the attack and risk surface, while NIST SP 800-53 Rev 5, OWASP ASVS and NIST CSF 2.0 set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP Agentic AI Top 10 | ASI03 — Identity & Privilege Abuse | Project instructions can cause an agent to overstep its intended authority. |
| ASI02 — Tool Misuse | Injected instructions can redirect a coding agent toward unsafe tool use. | |
| ASI09 — Human-Agent Trust Exploitation | These instructions exploit trust in repo-local guidance to shape agent behavior. | |
| Recommendation — Restrict agent actions so repository text cannot expand privilege or bypass approval. Constrain tool invocation paths and require approval for sensitive actions. Treat repository instructions as untrusted input unless explicitly validated. | ||
| NIST SP 800-53 Rev 5 | AC-6 — Least Privilege | Agentic tools should not gain broader access from local instructions alone. |
| SI-10 — Information Input Validation | Repo-scoped instructions are untrusted input that can alter downstream behavior. | |
| Recommendation — Limit agent permissions to the minimum needed for the current task. Validate and constrain instruction sources before they influence execution. | ||
| OWASP ASVS | V15 — Secure Coding and Architecture | Agent behavior and instruction trust boundaries are an architecture concern. |
| Recommendation — Design coding workflows so untrusted repository text cannot become operational policy. | ||
| MITRE ATT&CK | T1204 — User Execution | Malicious instructions can trick a user or assistant into performing unsafe actions. |
| T1059 — Command and Scripting Interpreter | The risk includes steering the agent into unsafe shell or script execution. | |
| T1552 — Unsecured Credentials | A key failure mode is instruction-driven secret discovery or theft. | |
| Recommendation — Hunt for execution paths where text guidance causes unsafe user or agent actions. Monitor and restrict scripted command execution initiated from agent workflows. Protect credentials from agent-accessible contexts and rotate exposed secrets quickly. | ||
| NIST CSF 2.0 | PR.AA-05 — Identity Management, Authentication, and Access Control | Access decisions should not be overridden by repository-local instructions. |
| Recommendation — Enforce access controls that remain independent of untrusted project content. | ||
Practitioner Guidance
What to verify: Treat any file that can steer an agentic coding tool as a policy artifact and confirm who can edit it, when it is loaded, and whether it can affect command execution or credential handling.
Common mistake: Teams often review these files for style or convenience, but not for authority. That misses the real issue, which is whether untrusted repository content can alter tool behavior without an explicit approval step.
Decision rule: If an instruction can change what the assistant may run, read, or disclose, place it behind review and approval controls rather than letting it inherit trust from the repo location alone.
Practitioner takeaway: The key security question is not whether the instruction is visible, but whether the agent treats it as binding. Once that happens, the file stops being documentation and starts behaving like an attack surface.
Related resources from NHI Mgmt Group
- Why do AI agents create more IAM risk than ordinary developer tools?
- Why do hidden skill fields create governance risk for agentic coding tools?
- Why do agentic coding tools create a different risk profile from standard developer tools?
- How should security teams handle project-scoped configuration that can spawn arbitrary processes in agentic coding tools?