TL;DR: A few lines in Claude Code’s CLAUDE.md file can override safety guardrails, trigger credential theft, and turn a developer assistant into an attack tool without coding skills, according to LayerX Security. The finding exposes a trust model that assumes project instructions are benign, even when they can redirect an autonomous coding assistant into harmful action.
Editorial analysis by NHI Mgmt Group, based on content published by LayerX Security: “Vibe Hacking: Claude Code Can Be Turned Into A Nation-State-Level Attack Tool With No Coding At All”.
Key questions
Q: What breaks when project prompt files can change an AI coding assistant's behaviour?
A: The trust boundary breaks down.
Q: Why do project-scoped instructions create risk for agentic coding tools?
A: They create risk because the assistant may treat those instructions as inherited authorisation, not as untrusted text.
Q: How can security teams detect whether an AI agent is inheriting unsafe instructions?
A: Look for files that shape tool use, execution scope, or approval logic, then review whether those files can be changed by ordinary contributors.
Practitioner guidance
- Treat project instruction files as controlled policy Classify CLAUDE.md and similar files as security-relevant assets, require review for any change that affects tool use or permissions, and include them in code review and access governance.
- Separate developer guidance from enforcement Keep convenience instructions distinct from rules that govern what an agent may execute, especially where local files can be modified by contributors with ordinary repository access.
- Limit repository write access on agent-enabled projects Review who can modify project-scoped prompt files, apply least privilege to contributor roles, and treat those edits as a path to behavioural change in the assistant.
Bottom line: Malicious project instructions can redirect an agentic coding assistant into harmful actions even when the same request would be rejected in a direct chat flow.
Explore further
View Full Forum → | NHI Foundation Course → | Our Services → | Read the full analysis →
Project instructions are an identity boundary, not documentation: CLAUDE.md was designed for developer guidance, but this case shows that the file can function as standing authorization for an agentic executor. That assumption fails when the actor can take actions on the local machine and treat repository text as policy. The implication is that identity programmes must classify project instruction files as governance artefacts that shape runtime authority.
A few things that frame the scale:
- 1 in 4 organisations are already investing in dedicated NHI security capabilities, with an additional 60% planning to do so within the next twelve months, according to The State of Non-Human Identity Security.
- Lack of credential rotation is cited as the top cause of NHI-related attacks by 45% of organisations, followed by inadequate monitoring and logging at 37%, according to the same research.
A question worth separating out:
Q: Who is accountable when an AI assistant follows malicious repository instructions?
A: Accountability sits with the organisation that allowed mutable instructions to act as standing authority without governance. If a project file can change agent behaviour, then ownership of that file, its review process, and its execution scope must be defined. Without that, the organisation has delegated security decisions to uncontrolled context.
👉 Read our full editorial: Claude Code trust assumptions collapse under malicious project prompts
Project instruction files are now part of the authorisation boundary: CLAUDE.md is not documentation when the assistant treats it as operating context. It becomes a policy input that can expand or redirect runtime behaviour, which means repository write access can translate into indirect control over an AI agent. The implication is that identity governance must cover instruction artefacts as controlled assets, not merely the accounts that use them.
A few things that frame the scale:
- Claude Code-assisted commits leaked secrets at a rate of 3.2%, more than double the human-only baseline of 1.5%, with peaks reaching 31 secrets per 1,000 commits in August 2025, according to the State of Secrets Sprawl 2026.
A question worth separating out:
Q: Should organisations treat AI project prompt files like code or documentation?
A: They should treat them like code or policy, not passive documentation. If a file can change execution behaviour, permissions, or safety decisions, it belongs in the same governance path as other assets that can materially affect system behaviour.
👉 Read our full editorial: Claude Code trust assumptions collapse under malicious project prompts