Coding agents can turn clues into actions because they optimize for solving the task, not for preserving a human’s trust boundary. When an attacker hides steps inside a puzzle or workflow, the model may treat those steps as legitimate progress and quietly execute them. That makes indirect prompt injection dangerous, especially when tool access includes file reads, network calls, or submission actions.
Why clues are more dangerous than commands in coding agents
Coding agents do not only execute explicit prompts, they also interpret surrounding context as part of the task. That matters because an attacker can hide malicious steps inside documentation, issue text, repository files, build output, or a workflow that looks like normal progress. The agent may treat those clues as legitimate evidence and carry them into tool use.
The security issue is not simple obedience, it is delegated execution under weak trust boundaries. If the agent can read files, call networked services, or submit changes, then the attacker does not need to ask for exfiltration directly. They only need to shape the context so the agent believes exfiltration is an efficient path to solving the task.
How indirect prompt injection turns context into exfiltration
Indirect prompt injection works when untrusted content is converted into instructions by the model’s reasoning loop. A poisoned README, issue comment, code snippet, or artifact can embed steps that look like clues, such as “verify this token,” “fetch the referenced file,” or “confirm the build status.” If the agent has broad tool access, those steps can lead to secret reads, outbound requests, or data submission.
This is why the risk grows with capability. A passive model can read a clue without consequence, but an agent that can browse, shell out, or edit files can turn the same clue into action. The danger is especially high when the agent’s tool permissions are wider than the current task actually requires.
Another problem is that clue-following feels legitimate to the system. The agent is usually optimising for task completion, not for preserving a human’s boundary between trusted instructions and untrusted content. That makes exfiltration paths easier to hide inside normal-looking troubleshooting, dependency resolution, or verification steps.
Why the risk scales with tool scope and authority
The more authority a coding agent has, the more ways a malicious clue can be operationalised. File-read access can expose local secrets, network access can send them out, and write access can alter code or configs so the compromise persists. When those permissions are combined in one workflow, the attack path becomes shorter and harder to notice.
That is why the best analogue is not “the agent was tricked into chatting,” but “the agent was tricked into acting.” In practice, the difference between a harmless clue and a harmful one is often whether the agent can move from interpretation to execution without a separate approval step.
Risk and Threat Considerations
Indirect prompt injection becomes most dangerous when the agent can cross from reading untrusted context to using privileged tools. The attacker’s objective is usually to smuggle instructions into normal work so the agent voluntarily reaches for secrets, makes outbound calls, or performs a submission on the attacker’s behalf.
Failure mechanism: Untrusted repository or workflow content is treated as task-relevant evidence, then converted into tool actions that exceed the user’s intended trust boundary.
Impact: Secrets, source code, tokens, build artifacts, or other sensitive data can be exposed, and the resulting exfiltration may look like ordinary agent activity unless it is separately logged and reviewed.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Agentic AI Top 10 and OWASP Non-Human Identity Top 10 define the specific risk controls and attack patterns relevant to this topic.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP Agentic AI Top 10 | ASI02 — Tool Misuse | Hidden clues become harmful when an agent misuses tools to act on them. |
| ASI03 — Identity & Privilege Abuse | Exfiltration risk increases when the agent can act with broader privilege than the task needs. | |
| ASI09 — Human-Agent Trust Exploitation | Indirect prompt injection exploits the trust relationship between user intent and agent execution. | |
| Recommendation — Restrict tool actions so untrusted context cannot drive unauthorized reads or exfiltration. Apply least privilege and per-action authorization to agent identities. Add approval gates where user intent and untrusted content could be confused. | ||
| OWASP Non-Human Identity Top 10 | NHI-02 — Secret Leakage | The scenario centers on secrets being exposed through agent-driven actions. |
| NHI-05 — Overprivileged NHI | Coding agents with broad tool access can turn cues into high-impact actions. | |
| NHI-10 — Human Use of NHI | The trust boundary breaks when human intent and agent execution are blended. | |
| Recommendation — Limit secret exposure to agent contexts and rotate any exposed credentials quickly. Reduce agent privilege to the minimum access needed for the current task. Separate human requests from agent execution paths and require explicit approvals. | ||
Practitioner Guidance
What to prioritise: Treat tool scope as the main control point. A coding agent that can only read the minimum files it needs is far less exposed than one that can freely browse, shell out, and submit results across a workspace or network boundary.
What to verify: Confirm that untrusted content cannot directly trigger high-impact actions without an independent approval checkpoint. The key test is whether a hidden instruction in a file, issue, or webpage could cause the agent to read secrets or make an external call before a human sees the request.
Common mistake: Relying on “the model should know better” is not a control. If the agent has broad authority, the system is assuming perfect judgment from a component that is explicitly designed to follow cues and optimise completion.
Practitioner takeaway: The exfiltration risk rises when the agent can translate clues into actions, so the practical defence is to narrow authority, isolate untrusted context, and require separate checks before any action that can expose data.
Related resources from NHI Mgmt Group
- Why do AI coding agents create governance risk even when they improve productivity?
- Why do AI coding agents create security risk even when they use the same model?
- Why do AI coding agents create access and governance risk even when they are not autonomous?
- Why do AI agents create higher risk when they can access payment records and refund tools?
Deepen Your Knowledge
Reviewed and updated by the NHIMG editorial team on September 30, 2026.
NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org