TL;DR: Years of ProdSec review history have been turned into 343 rules across 16 vulnerability categories by a SAGE pipeline using a multi-model Finder, Critic, and Judge, while cutting review time and hardening against prompt injection, according to 1Password. The core lesson is that AI-assisted security review only works when discovery, verification, and adjudication are separated.
At a glance
What this is: 1Password describes how it scaled product security review by splitting AI-assisted code analysis into separate recall, critique, and verdict stages backed by 343 rules across 16 vulnerability categories.
Why it matters: This matters because security teams using AI in review workflows need a governance model that preserves high recall without letting unverified findings or prompt injection drive decisions.
By the numbers:
- The team distilled 8,343 reviews from nearly 9,000 pull requests into its ruleset.
- SAGE v1 now has 343 rules across 16 vulnerability categories.
- The average token cost per scan was $0.47 USD in the v0 implementation.
Context
Security code review breaks down when the volume of pull requests grows faster than human reviewers can adjudicate findings consistently. In AI-assisted review, the governance problem is not just detection quality, but how to preserve evidence, challenge findings, and avoid turning a noisy model output into a trusted decision.
This article sits at the intersection of developer security, AI-assisted analysis, and security operations workflow design. The primary identity question is not human authentication or machine credentials, but how a security control behaves when an AI system is inserted into a judgment-heavy engineering process.
1Password's example is a production-style pattern for security review at scale, not a generic AI demo. The article shows why separating pattern discovery from verification matters once code review volume and model-assisted development both rise.
Key questions
Q: How should security teams use AI-assisted code review safely?
A: Use it as a triage layer that accelerates first-pass detection, then require a separate validation step for findings that affect access control, authentication, secrets, or release gating. The safest pattern is hybrid review, where deterministic analysis and human judgement backstop the model’s reasoning.
Q: Why do AI review pipelines need to treat prompt injection as a security issue?
A: Because the code or context being analysed is untrusted input that can steer the model away from correct judgment. If the review system cannot detect manipulation attempts, attackers can influence the control itself, not just the code under review.
Q: What are the signs that an AI security review workflow is over-reliant on one model?
A: A strong signal is when discovery, verification, and final approval all come from the same prompt or provider, with no independent challenge stage. Another sign is that the workflow breaks when the model changes, because the control logic and evidence trail were never separated from the vendor choice.
Q: What should teams do when AI review flags look noisy but useful?
A: Preserve the noisy output for recall, then measure which findings survive adversarial critique and final adjudication. The right question is not whether the first pass is perfect, but whether the pipeline can convert broad detection into defensible security decisions.
Technical breakdown
Why AI-assisted code review needs separate recall and proof stages
AI review systems fail when one model is asked to both spot issues and prove them. High recall is useful for surfacing suspicious patterns, but proof requires different inputs, different incentives, and often a different model. In SAGE, the Finder stage is intentionally broad and noisy, the Critic stage pressure-tests each finding using the rule body and code hunks, and the Judge stage makes the final call. That separation reduces false confidence and makes the workflow auditable. The key design principle is progressive disclosure, where later stages see only what they need to decide.
Practical implication: split detection from adjudication so security reviewers can inspect why a finding survived critique, not just whether a model flagged it.
How prompt injection changes the security posture of AI review pipelines
Once an LLM is used inside a security workflow, attacker-supplied text can become an input to governance decisions. In a code review context, prompt injection means malformed or malicious content in code or surrounding context tries to steer the model away from correct analysis. SAGE treats prompt-injection attempts as findings in their own right, which is important because the review tool itself becomes part of the attack surface. A pipeline that only looks for code defects but ignores model manipulation can be bypassed by adversarial content embedded in the material under review.
Practical implication: detect and log prompt-injection attempts separately from code findings so model manipulation is visible to security operations.
What model and vendor agnosticism adds to operational resilience
A provider-agnostic interface matters because AI-assisted security review will not stay stable around one model vendor or one capability tier. SAGE uses a common client layer so the Finder, Critic, and Judge can be reassigned as model quality, cost, and latency change. That reduces vendor lock-in and lets the control evolve without rewriting the workflow. It also supports a more defensible governance model: the system's assurance comes from process separation and evidence handling, not from trusting a single model to be consistently right.
Practical implication: build the review harness so the underlying models can change without altering the control logic or evidence trail.
Breaches seen in the wild
- reviewdog Action compromise 2025: A stolen maintainer token poisoned reviewdog/action-setup, leaking CI secrets including the tj-actions bot token used in the next attack.
- CI/CD pipeline exploitation case study: Credentials in an exposed .git/config let a researcher edit a Bitbucket pipeline so it planted their SSH key on the server. No victim was named.
Read and download The State of NHI & AI Agent Breach Report 2026, covering 200+ breaches impacting Non-Human Identities including AI Agents.
NHI Mgmt Group analysis
AI-assisted security review is a governance workflow, not a single model prompt. The article shows why discovery, challenge, and verdict need separate control points once a model starts influencing security outcomes. That matters because the assurance question is no longer whether the model can find issues, but whether the surrounding process can prove and explain why a finding stands.
Progressive disclosure is the right control pattern for high-volume security review. A broad finder is useful only when a later stage can challenge it with stronger evidence and a neutral verdict step can resolve conflicts. This is a workflow design problem, not a model-selection problem, and practitioners should treat it that way.
Prompt injection is part of the review threat model, not an edge case. Any security control that reads untrusted code, comments, or repository content is exposed to manipulation attempts. The practical implication is that AI review systems need their own defensive logic and logging, because the analysis layer itself can be targeted.
Model portability is now a control requirement, not a convenience feature. When security review depends on a single AI provider, the governance model inherits that provider's blind spots, pricing, and release cadence. A detachable harness lets teams preserve control continuity while swapping models as their risk profile changes.
Separated recall and proof creates the right kind of audit trail. The most valuable output is not the first model's suspicion, but the evidence path that survives criticism and is recorded for human review. That is how AI review becomes governable at scale instead of merely faster.
From our research library:
- Only 44% of developers are reported to follow security best practices for secrets management, exposing a significant developer behaviour gap, according to the State of Secrets in AppSec.
- Read next: Agentic AI Security Guide
What this signals
Separated recall and proof is the emerging control pattern for AI-assisted security review. Security teams should expect the governance boundary to move from the model prompt to the workflow around it. The strongest systems will be the ones that can explain why a finding survived critique, not merely that it was detected.
AI review controls now need their own failure modes. If a pipeline reads untrusted code, comments, or repository metadata, prompt injection becomes an operational risk, not a theoretical one. That means detection, logging, and human review need to be designed for model manipulation from the outset.
Developer behaviour still matters, even in an AI-assisted workflow: only 44% of developers are reported to follow security best practices for secrets management, according to the State of Secrets in AppSec. The practical signal is that automation can scale review, but it cannot compensate for weak secure-coding habits.
For practitioners
- Separate finding generation from final adjudication Use different stages, prompts, and evidence inputs for recall, critique, and verdict so the same model is not asked to discover and prove the same issue.
- Treat prompt injection as a security signal Log injection attempts as their own findings and route them to the same review queue as code defects so model manipulation is observable.
- Preserve a machine-readable ruleset Keep a compact index of rules and summaries so each finding can be traced back to a specific control, category, and body of guidance.
- Design for provider switching Abstract the model interface so evaluation quality, latency, and cost can change without rewriting the security workflow or evidence trail.
- Keep a human review backstop Retain analyst oversight for edge cases, false positives, and unresolved conflicts, especially where code context is incomplete or ambiguous.
Key takeaways
- AI-assisted code review becomes governable only when discovery, challenge, and verdict are separated into distinct stages.
- Prompt injection is part of the security control surface when an AI system reads untrusted code or repository context.
- A provider-agnostic harness gives security teams more durable control than a workflow tied to one model vendor.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Agentic AI Top 10 and OWASP API Security Top 10 address the attack and risk surface, while NIST AI RMF and NIST CSF 2.0 set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP Agentic AI Top 10 | ASI02 — Tool Misuse | The article centres on a multi-stage AI workflow whose outputs must be constrained and verified. |
| ASI06 — Memory & Context Poisoning | Prompt injection and untrusted repository content are direct context-poisoning concerns. | |
| Recommendation — Constrain model stages so discovery, critique, and verdict each operate within narrow, auditable responsibilities. Treat repository text as hostile input and isolate it from final decision logic wherever possible. | ||
| NIST AI RMF | GOVERN — AI Governance and Accountability | The article is fundamentally about governing an AI-assisted security decision process. |
| Recommendation — Define accountable review ownership, evidence handling, and escalation paths for AI-assisted security decisions. | ||
| NIST CSF 2.0 | PR.AA-05 — Access Permissions, Entitlements and Authorizations | The workflow governs who and what may influence security review decisions and outputs. |
| Recommendation — Limit AI review influence to authorised workflow stages and preserve human approval at decision points. | ||
| OWASP API Security Top 10 | API10 — Unsafe Consumption of APIs | The pipeline consumes model APIs as part of a security control and must handle them safely. |
| Recommendation — Validate model outputs before consuming them as control inputs in downstream security workflows. | ||
Key terms
- Progressive Disclosure Pipeline: A review design that reveals information in stages instead of giving one system full context at once. In security operations, it reduces overreach and makes each decision step easier to audit. For AI review, it also limits prompt injection because later stages see only the evidence they need.
- Prompt Injection (Agentic): An attack where malicious instructions are embedded in content that an AI agent reads, causing the agent to execute unintended actions using its own legitimate credentials. A primary vector for agent goal hijacking and identity abuse.
- Model Agnosticism: A design approach that avoids dependence on one AI model by making the workflow resilient to model swaps, capability changes, and provider constraints. In security terms, the important question is not which model is best, but whether controls still hold when the contributor changes.
- False Positive Filter: A control that removes speculative findings before they reach final approval or remediation. In AI-assisted security review, it prevents noisy outputs from overwhelming engineers and keeps the approval queue focused on credible issues. It is only effective when the filter is independent from the initial detector.
Deepen your knowledge
NHI governance, agentic AI identity, and machine identity lifecycle are core topics in our NHI Foundation Level course, the industry's only accredited NHI security programme. If you are building or maturing an IAM programme, it is worth exploring.
Published by the NHIMG editorial team on July 1, 2026.
Updated on October 11, 2026.
NHI Mgmt Group, the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org