The trust boundary breaks down. A repository file can quietly alter what the assistant believes it is allowed to do, so security controls that focus only on user prompts miss a persistent policy layer that can authorise harmful behaviour across every future session.
What breaks in the trust model when a repository file can steer an AI coding assistant?
The trust boundary no longer sits at the chat box. A checked-in instruction file becomes persistent policy, so the assistant is no longer governed only by the current prompt or the visible user session. That matters because repository content can outlive one conversation, influence future runs, and shape actions in ways the operator may not notice.
That shift is especially dangerous in agentic workflows, where the assistant can execute commands, touch files, or invoke tools. The file is not just text, it becomes a control surface for behaviour, and once the assistant treats it as authoritative, the usual assumption that each session starts clean no longer holds.
Why this is a trust-boundary failure, not just a prompt-injection problem
The core failure is persistence. A user prompt is transient, but a project prompt file can remain in the repository, travel with the codebase, and be reused by the assistant every time the project is opened. That makes it closer to policy than conversation, which means hidden or malicious instructions can compound across sessions and across developers.
For an ai coding assistant, that changes the security question from “what did the user ask now?” to “what instructions is the tool carrying forward from the project state?” The practical result is that controls focused only on immediate prompt content miss a durable layer of influence that can reframe allowed actions, suppress warnings, or encourage unsafe tool use.
It also blurs ownership. If a repository file can shape behaviour, then code review, branch protection, and change control become part of assistant security, not just software engineering hygiene. In a project with shared contributions, the highest-risk issue is often not obvious malice, but low-friction drift where a seemingly helpful instruction file accumulates broader permissions over time.
What security teams should inspect first in these files
Start by treating the instruction file as a policy artifact with real blast radius. If it can affect commits, shell commands, dependency installation, test execution, or access to secrets, it deserves the same review discipline as other high-impact automation inputs. A file that only tunes style is one thing; a file that steers tool invocation or environment assumptions is materially different.
Pay attention to instructions that normalise dangerous behaviour, such as suppressing confirmation, widening scope to the whole workspace, or trusting repository content more than the operator. Those are the patterns that convert a convenience feature into an execution policy. A useful internal reference point is the AI Coding Agents Security Guide, which frames this as a combined prompt, secrets, and sandboxing problem rather than a simple chatbot issue.
Repository-scoped instructions also interact with overprivileged tokens and workspace trust. In practice, the boundary is weakest when the assistant can read the repo, reach external services, and act with the developer’s authority all at once. That is why the control problem is not just “sanitize prompts”, it is “constrain what persistent project instructions are allowed to influence”.
Risk and Threat Considerations
Persistent project prompt files create a durable abuse path because an attacker only needs to plant or modify one repository artifact to influence many future assistant sessions. If that file is trusted as part of the working context, malicious instructions can steer code generation, credential exposure, dependency installs, or destructive commands long after the initial change was merged.
Failure mechanism: A repository-controlled instruction layer is treated as legitimate policy, so the assistant follows attacker-supplied or overbroad guidance across repeated sessions, including tool-enabled actions.
Impact: The result can be unauthorized code changes, secret exposure, data loss, supply-chain contamination, or repeated unsafe behaviour that looks like normal assistant output rather than a one-time prompt attack.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Agentic AI Top 10 and OWASP Non-Human Identity Top 10 address the attack and risk surface, while NIST SP 800-53 Rev 5 and NIST Zero Trust (SP 800-207) set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP Agentic AI Top 10 | ASI03 — Identity & Privilege Abuse | Persistent repo instructions can steer agent authority and tool use. |
| Recommendation — Constrain agent permissions and review any file that can alter authority or actions. | ||
| OWASP Non-Human Identity Top 10 | NHI-10 — Human Use of NHI | Project files can let humans indirectly control non-human assistant behaviour. |
| NHI-05 — Overprivileged NHI | AI assistants become risky when repository instructions expand their authority. | |
| NHI-01 — Improper Offboarding | Stale project instructions can keep influencing assistants after ownership changes. | |
| Recommendation — Separate human-authored guidance from machine-executed policy and review both. Audit assistant permissions and remove any excess authority. Remove obsolete instruction files and revoke outdated automation paths promptly. | ||
| NIST SP 800-53 Rev 5 | AC-6 — Least Privilege | Assistant tools should not inherit broad authority from repository content. |
| CM-5 — Access Restrictions for Change | Repository files that alter behaviour are change-controlled policy inputs. | |
| SI-10 — Information Input Validation | Persistent instructions are untrusted inputs that shape runtime behaviour. | |
| Recommendation — Limit assistant tool permissions to the minimum needed for the task. Restrict who can modify project instruction files and review those changes. Validate and constrain repository-derived instructions before the assistant acts. | ||
| NIST Zero Trust (SP 800-207) | Never Trust, Always Verify | Repository context should not be trusted as inherently safe policy. |
| Recommendation — Verify each instruction source before granting it influence over actions. | ||
Practitioner Guidance
What to prioritise: Classify every project instruction file by what it can influence. If it can change tool execution, file writes, command running, dependency retrieval, or secret handling, review it as a security-relevant control surface, not as documentation.
What to verify: Confirm whether the assistant treats the file as authoritative across sessions, whether contributors can edit it through normal workflow, and whether the file can affect privileged operations without a second approval step. If the answer is yes to any of those, the trust boundary is already wider than many teams assume.
Common mistake: Teams often harden the chat experience and leave the repository context untouched. That misses the persistent policy layer, which is the part most likely to survive one prompt and shape the next hundred.
Practitioner takeaway: The safest mental model is that repository instructions are executable policy, so their change control, scope, and authority must be governed like any other security-sensitive automation input.
Related resources from NHI Mgmt Group
- What breaks when an AI coding assistant is allowed to read files but not inspect data sensitivity?
- What breaks when an AI coding assistant executes project content before trust is confirmed?
- How should security teams govern AI agents that can change behaviour based on prompt context?
- What breaks when AI coding agents can read project setup metadata?