Because the repository can influence both what gets inspected and what gets executed. If a repo can select the reviewer model, narrow its scope, and then send the agent into its own tests, the security process becomes partly authored by the code under review. That creates a trust boundary problem, not a model failure, and it can hide malicious behavior in plain sight.
How repo-defined workflows change the trust boundary for AI coding agents
Repo-defined workflows are risky because the repository itself can influence both the reviewer’s scope and the actions it is asked to take. In practice, that means the untrusted code under review may shape the inspection path, the model choice, and even the follow-on execution path. The security issue is not just “bad code,” but code that can steer the review process around its own malicious behavior.
That is why repository-controlled instructions are qualitatively different from ordinary source code. A normal review assumes the reviewer owns the process; repo-defined workflows let the repo partially author it. Once the workflow can narrow what is visible, redirect the agent into tests, or change how tools are invoked, the trust boundary moves from “code being inspected” to “code influencing the inspector.”
For AI coding agents, that boundary shift matters because the agent is not only reading content, it is often executing instructions, calling tools, and handling credentials or environment state while it reviews. When the repository can participate in that orchestration, a malicious contributor can hide intent behind apparently routine automation and make the review itself part of the attack surface.
What makes this a supply-chain and review integrity problem
This pattern is best understood as a review-integrity problem inside the software supply chain. The repository can present safe-looking files to the agent while steering it toward other paths, or it can cause the agent to exercise code that behaves differently when it is “being reviewed” than when it is shipped. That creates asymmetry between what the reviewer thinks it is checking and what actually runs.
Repo-defined workflows also create room for scope manipulation. If the workflow can limit file selection, alter prompts, or decide when tests run, the agent may never inspect the most dangerous path at all. If the workflow can induce execution, the review environment can become a place where hidden commands, data access, or side effects occur under the appearance of validation.
That is why this class of issue is not just about the model’s judgment. It is about whether the review process itself is independently trustworthy when the subject under review can help define the process. In trusted build and review pipelines, that separation is essential.
Why untrusted repositories can turn ordinary automation into a hidden attack path
The practical risk is that the repository can exploit the agent’s helper role. A workflow file, test command, or repo-local instruction can look like routine automation while actually shaping the agent’s behavior toward attacker goals. In the worst case, a review agent becomes a trigger for execution, credential exposure, or false confidence, because the repository controls enough of the path to make the malicious activity look expected.
This is especially dangerous when review and execution happen in the same environment. If the agent can read secrets, reach internal services, or run repo tests without strong isolation, an attacker can use the review as a bridge from code inspection into operational impact. For that reason, repo-defined workflows should be treated as part of the untrusted input, not as neutral scaffolding around the review.
For a broader agentic AI lens, the key question is whether the repository can influence the agent’s authority, not just its output. NHIMG’s Agentic AI Security Guide and AI Coding Agents Security Guide both frame that distinction around tool access, execution boundaries, and blast radius.
Risk and Threat Considerations
Repo-defined workflows create a compromise path where malicious code can influence what the reviewer sees, what it runs, and which protections are bypassed. That makes the review process itself a target, especially when the agent has enough access to touch tests, commands, or environment state.
Failure mechanism: The repository supplies workflow instructions that narrow inspection, alter tool use, or cause execution in a way that conceals malicious behavior and extends attacker control into the review pipeline.
Impact: Reviewers can miss harmful code, accidentally execute attacker-chosen actions, or expose credentials and internal resources while believing they are performing a safe inspection.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Agentic AI Top 10 and MITRE ATT&CK address the attack and risk surface, while NIST SP 800-53 Rev 5 sets the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP Agentic AI Top 10 | ASI03 — Identity & Privilege Abuse | Repo workflows can steer agent authority and tool use. |
| ASI02 — Tool Misuse | The risk involves repo-driven commands and tests being misused. | |
| ASI01 — Agent Goal Hijack | A malicious repo can redirect the agent away from the real review goal. | |
| Recommendation — Constrain agent authority and require external approval for workflow-driven execution. Restrict tool execution paths to trusted commands and sandbox all repo-influenced actions. Validate that repository content cannot redefine the agent’s review objective or scope. | ||
| MITRE ATT&CK | T1218 — System Binary Proxy Execution | Workflow-controlled execution can proxy harmful actions through trusted tooling. |
| Recommendation — Monitor for trusted-tool execution patterns that proxy attacker-chosen commands. | ||
| NIST SP 800-53 Rev 5 | AC-6 — Least Privilege | Agents reviewing untrusted code should not have excess execution authority. |
| SA-11 — Developer Testing and Evaluation | The issue sits in the trustworthiness of review and test execution. | |
| Recommendation — Apply least privilege so repo-driven workflows cannot expand the agent’s access. Require independent validation of workflow-controlled tests before trusting results. | ||
Practitioner Guidance
What to verify: Treat repo-local workflow definitions as untrusted inputs and confirm that they cannot select privileged tools, expand execution scope, or alter reviewer policy without an external trust decision. If they can, the review path is already contaminated.
What good looks like: The agent’s model choice, prompt scope, tool permissions, and execution environment are fixed outside the repository, and any repo-triggered test or command runs in a bounded sandbox with no standing access to sensitive credentials.
Decision rule: If a workflow can affect both inspection and execution, review it as part of the threat surface before trusting any conclusions from the agent. If you cannot separate review control from repo control, do not use the workflow to validate untrusted code.
Practitioner takeaway: The core control is separation of authority, not smarter review. Untrusted code should never be able to shape the process that is supposed to judge it.