They collapse the boundary between instructions and data. Repository files, MCP output, and marketplace extensions can all be treated as trusted context before a human sees them. That makes a prompt, config file, or tool response capable of driving commands, privilege use, or data exposure unless a separate runtime layer inspects the action and blocks unsafe behavior in real time.
Why default trust turns AI coding agents into high-impact actors
AI coding agents become riskier when they treat repository content and tool output as trusted because the boundary between instruction and evidence disappears. A file, dependency, README, MCP response, or extension payload can quietly influence what the agent edits, runs, or exfiltrates. That means the agent is no longer just assisting a developer, it is executing decisions on inputs that may already be compromised.
The core failure is not “AI making mistakes” in the abstract. It is that the agent often receives the same operational weight for untrusted text as it does for a deliberate human instruction. Once that happens, the highest-risk action is not the model’s answer, but the commands, privilege use, and side effects the agent triggers before anyone reviews them.
Default trust also broadens the blast radius. In a normal code review flow, suspicious content can be seen, questioned, or blocked before execution. In an agentic flow, the agent may act first, so a poisoned repository, malicious extension, or deceptive tool response can move from suggestion to execution in a single turn. AI Coding Agents Security Guide covers the controls that matter when agent context is treated as an execution surface.
How repository content and tool output become an attack path
Repository files are dangerous when they can steer the agent through hidden instructions, configuration, or dependency choices. A poisoned README, agents.md file, CI hint, or template can cause the agent to run commands that look locally reasonable but are actually attacker-directed. Tool output is similarly risky because the agent may interpret returned text as ground truth, even when the tool result is fabricated, manipulated, or shaped by a prior compromise.
The most important pattern is indirect prompt injection. The attacker does not need to talk to the agent like a user does. They only need to place content where the agent will read it and then trust it. That is why repository trust and tool trust are especially dangerous in coding workflows: they create a path from untrusted text to action without an explicit human authorization step. Gemini CLI prompt injection flaw 2025 shows how a poisoned README can turn into command execution and secret exposure, while Amazon Q MCP config vulnerability 2026 demonstrates how a malicious repository can hijack an agent through its own configuration.
Tool output creates a second trust problem: the agent may use the output to justify a next action, then chain that action into broader access or data movement. When the tool is connected to credentials, cloud APIs, source control, or production systems, the output is not just information, it is an instruction vector that can drive real privilege use. Sentry MCP Agentjacking 2026 is a clear example of fake tool output leading to attacker-controlled behavior.
What practitioners need to separate before they let an agent act
The practical control point is to separate interpretation from execution. The agent may inspect repository content and tool output, but it should not be able to convert either into commands, writes, credential use, or external calls without a policy decision layer. That layer needs to understand the action, the target, the identity in use, and the blast radius, not just the text that triggered the action.
AI Agent Authorisation Guide is directly relevant here because the right model is per-action authorization, not blanket trust in the agent session. In practice, the useful question is whether the next step is still safe if the repository, extension, or tool response is hostile. If the answer is no, the action needs approval, scope reduction, or a hard block.
At scale, the highest-value checks are task scoping, environment separation, and runtime enforcement. Agents that can read everything but act on very little are safer than agents that can freely chain read, decide, and execute. Zero Trust for AI Agents is the right lens for this because the agent, its context, and its tools should all be verified before each sensitive action.
Risk and Threat Considerations
Default trust creates a direct attack path from untrusted content to privileged action. If the agent can read repository files or tool responses and immediately use them as operating context, an attacker only needs one poisoned file, one misleading tool response, or one compromised extension to influence code changes, credential use, or destructive commands. That is especially dangerous when the agent has access to developer tokens, cloud APIs, or production-adjacent systems.
Failure mechanism: The agent collapses instruction, evidence, and execution into one trust domain, so malicious text can be treated as an approved directive before any separate policy check or human review occurs.
Impact: The result can be unauthorized code changes, secret exposure, supply-chain contamination, cloud actions, or irreversible data loss, often with the agent’s own credentials and audit trail making the compromise look legitimate.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Agentic AI Top 10, OWASP Non-Human Identity Top 10 and OWASP API Security Top 10 address the attack and risk surface, while NIST SP 800-53 Rev 5 and NIST Zero Trust (SP 800-207) set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP Agentic AI Top 10 | ASI03 — Identity & Privilege Abuse | Default trust lets agents misuse inherited privileges from untrusted inputs. |
| ASI02 — Tool Misuse | Repository and tool output can steer an agent into unsafe tool actions. | |
| Recommendation — Enforce per-action authorization and block agent privilege use without policy checks. Restrict tool invocation to approved actions and validate tool-driven requests at runtime. | ||
| OWASP Non-Human Identity Top 10 | NHI-04 — Insecure Authentication | Trusted context can push agents to use exposed credentials or tokens unsafely. |
| NHI-05 — Overprivileged NHI | Agent risk rises when default trust meets excessive permissions and broad access. | |
| Recommendation — Separate agent context from credential use and require stronger runtime verification. Reduce agent permissions to the minimum task scope before allowing execution. | ||
| NIST SP 800-53 Rev 5 | AC-6 — Least Privilege | Agents should not inherit broad execution power from untrusted repository or tool content. |
| IA-5 — Authenticator Management | Tool-driven attacks often aim to expose or misuse credentials and tokens. | |
| AU-6 — Audit Record Review, Analysis, and Reporting | Runtime inspection and attribution are needed when agents act on untrusted inputs. | |
| Recommendation — Limit agent actions to the minimum privileges needed for the task. Protect, rotate, and tightly govern the credentials an agent can access. Log agent actions and review anomalous command or tool-driven behavior. | ||
| NIST Zero Trust (SP 800-207) | 3.1 — Never trust, always verify | The question is fundamentally about removing default trust from agent execution. |
| Recommendation — Verify each request and action instead of trusting context by default. | ||
| OWASP API Security Top 10 | API5 — Broken Function Level Authorization | Agents that call tools or APIs need action-level authorization, not assumed access. |
| Recommendation — Authorize each agent function call before allowing execution. | ||
Practitioner Guidance
What to verify: Verify that repository content and tool output are only inputs, not automatic authority. The agent should not be able to escalate from “I saw this text” to “I executed this action” without a separate control point that can inspect the request and the identity in use.
Decision rule: If the next action can modify code, run a command, or touch secrets, require explicit policy approval or a human step unless the environment is tightly sandboxed and the scope is minimal.
What good looks like: The agent can suggest, summarize, and draft, but sensitive actions are bounded, attributable, and reversible. The safe state is not zero autonomy, it is autonomy with clear execution boundaries.
Practitioner takeaway: Treat every repository file and tool response as potentially hostile until a runtime control has validated the action, because default trust is what turns simple context into executable compromise.