By NHI Mgmt Group Editorial TeamBased on Pillar Security: “Untrusted Project-Local Filters in RTK: When Your AI's Eyes Are Someone Else's to Control” (May 20, 2026)

TL;DR: RTK’s project-local filters let repository content control what Claude Code could see, creating a medium-severity path to hide backdoors and suppress scanner output before review, according to Pillar Security. The deeper issue is trust laundering: if the observation layer is attacker-shaped, AI-assisted code review cannot be treated as a reliable control.


At a glance

What this is: This research shows that RTK's repository-local filters could be used to hide malicious code and security warnings from Claude Code before review.

Why it matters: IAM and security teams should treat information-preprocessing layers around AI coding assistants as part of the trust boundary, because attacker-shaped observation can undermine both human and machine review.


Context

RTK is a preprocessing layer for AI coding assistants that filters command output before the model sees it. The security problem is not execution control alone, but observation control: if untrusted repository content can shape what the assistant reads, then review results can be silently distorted.

In this case, project-local filters were loaded automatically from cloned repositories with no approval prompt. That creates a governance gap for AI-assisted development, because the system trusted repository-originated configuration as if it had the same authority as user-controlled settings.

The article's central claim is that tools sitting between code and the model are part of the attack surface. For AI-assisted code review, trust in the pipeline matters as much as trust in the model output itself.


Key questions

Q: What breaks when project-local AI filters load automatically from a repository?

A: Automatic loading breaks the assumption that the reviewer sees the same evidence the repository contains. An attacker can hide malicious lines from diffs, security scanner output, or source files before the model sees them. The result is false confidence in a clean review and a path for compromised code to advance toward production.

Q: Why do attacker-controlled preprocessing layers create risk even when they never execute code?

A: Because the risk is perceptual, not just executable. If a tool can suppress warnings, trim diffs, or hide lines from the model, it can change the assistant's conclusion without touching runtime behaviour. That makes observation-layer controls part of the security boundary for AI-assisted development, especially in shared repositories.

Q: How do security teams know whether AI review outputs are actually trustworthy?

A: Teams need to validate the integrity of the entire observation chain, from repository files to the model’s context window. If any preprocessing layer can remove evidence without review or provenance checks, a clean answer may only mean the input was filtered. Compare AI output against raw artefacts where possible.

Q: Should organisations treat filters, hooks, and plugins as supply-chain controls?

A: Yes. Anything that preprocesses code, logs, or scan output before AI review should be governed like a supply-chain control because it can suppress evidence or alter context. That means reviewing origin, enforcing change detection, and separating user-controlled settings from repository-controlled settings.


Technical breakdown

How project-local filters change the model's observation layer

RTK sits between the shell output and the AI coding assistant, stripping or reshaping lines before they reach the model. That makes the filter layer part of the assistant's perception system, not just a convenience feature. If the filter configuration comes from the repository, an attacker can decide which evidence survives into context and which evidence disappears. The vulnerability is especially dangerous because the model does not see an obviously malformed prompt. It sees a normal-looking, already filtered result, which makes the manipulation difficult to detect downstream.

Practical implication: Treat preprocessing and filtering as governed inputs, not neutral tooling.

Why repository-supplied configuration becomes trusted authority

The core mistake is collapsing origin and content. A file like .rtk/filters.toml can look like ordinary configuration, yet if it is loaded automatically from a cloned repository it inherits authority from the tool rather than from the user. That is a trust-boundary failure, not a parsing bug. The attacker does not need execution rights inside the model. They only need commit rights, a compromised contributor account, or access to a repository where the project-local file is accepted automatically.

Practical implication: Require explicit trust decisions for repository-local configuration before it changes model-visible output.

How attacker-controlled filters suppress security review

Because RTK's filters are regex-driven, they can remove lines from command output, scanner findings, and git diffs alike. That means the same control can hide a backdoor in source code and erase the warning that would have exposed it. This is why the issue maps to a trust-laundering pattern in AI-assisted development: untrusted content is promoted into a trusted perceptual channel. Once the evidence is gone from the assistant's view, the review workflow can produce a false clean result with no obvious failure signal.

Practical implication: Assume any output-transforming layer can be used to launder evidence out of review.


Threat narrative

Attacker objective: To hide malicious code and security findings from AI-assisted review so compromised code ships without detection.

  1. An attacker commits a malicious .rtk/filters.toml file into a repository that a developer later clones.
  2. When the developer works with Claude Code and RTK, the project-local filter file is loaded automatically and alters what the assistant can observe.
  3. The attacker suppresses malicious lines, scanner warnings, or diff evidence so the assistant reviews an incomplete picture of the codebase.
  4. The downstream impact is that compromised code can pass AI-assisted review and automated scanning, then move toward production as if it were clean.

Read and download The State of NHI & AI Agent Breach Report 2026, covering 200+ breaches impacting Non-Human Identities including AI Agents.


NHI Mgmt Group analysis

Trust laundering is the real failure mode here: the repository was allowed to supply perceptual authority, not just configuration. That assumption was designed for content that was safe to load automatically because it only changed local behaviour. That assumption fails when the content can hide evidence from an AI reviewer, because the reviewer now depends on the very file it is supposed to inspect. The implication is that AI-assisted review needs a stricter trust boundary around observation than around execution.

Observation-layer controls need the same governance rigor as execution-layer controls: teams often review hooks, sandboxes, and approval policies while ignoring filters and preprocessors. This article shows that a non-executing tool can still determine what the model knows, which makes it security-relevant even if it never runs code. The practitioner takeaway is to classify output-shaping tools as part of the software supply chain, not as harmless ergonomics.

Project-local defaults are especially dangerous in collaborative codebases: a file committed by a contributor can silently override user expectations at clone time. That matters because AI-assisted development increasingly uses the same repository for code generation and code review. When both workflows inherit the same filtered view, the control fails in both directions at once, which turns a single trust mistake into a pipeline-wide blind spot.

Hash-based trust revocation is the right design pattern for this class of problem: the fix described in the article ties trust to the reviewed file contents rather than to the file path or feature name. That is the correct conceptual model for repository-originated AI tooling controls. Practitioners should read this as evidence that trust decisions need to be content-bound, change-sensitive, and explicit.

Named concept: perceptual authority laundering: this article describes a class of failure where attacker-controlled configuration changes what an AI is permitted to observe, not what it is allowed to execute. That distinction matters because it widens the attack surface beyond prompts and code to any layer that mediates visibility. Security teams should treat AI review integrity as a pipeline problem, not a model problem.

From our research library:

What this signals

Perceptual authority is now a governance problem: AI-assisted development tools do not only need safe execution paths. They also need trusted observation paths, because an attacker who controls what the model can see can steer review outcomes without triggering traditional runtime alarms.

Repository-local defaults should no longer be treated as neutral: once a cloned project can supply configuration that shapes model output, the trust model must distinguish between user-owned settings and attacker-supplied settings. That distinction belongs in engineering policy, not just in tool documentation.

Access reviews are the wrong mental model for this class of control: the issue is not whether a privilege exists long enough to be reviewed, but whether the evidence ever reaches the reviewer intact. In AI pipelines, integrity of the observation path is the control point that matters.


For practitioners

  • Audit output-shaping tools in AI review pipelines Inventory every filter, hook, plugin, and preprocessor that can change what an AI assistant sees before it reviews code or scanner output.
  • Require explicit trust for repository-local settings Block automatic loading of project-supplied configuration until a reviewer has validated origin, content, and intended scope of influence.
  • Separate review evidence from repository content Preserve a raw, unfiltered path for security findings, diffs, and command output so the assistant cannot be shown a curated subset only.
  • Revoke trust when configuration changes Tie trust state to a hash of the reviewed file and invalidate it after any repository update, pull, or rewrite.

Key takeaways

  • RTK's project-local filters created a trust gap because repository content could shape what an AI assistant was allowed to see during code review.
  • The article ties that gap to a real vulnerability, CVE-2026-45792, and notes that RTK had more than 50,000 GitHub stars in five months.
  • The control failure is not the model itself but the observation layer, which means teams need explicit trust boundaries around preprocessors, filters, and repository-supplied configuration.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

OWASP Non-Human Identity Top 10, OWASP Agentic AI Top 10, MITRE ATT&CK and OWASP API Security Top 10 address the attack and risk surface, while NIST CSF 2.0 sets the governance and control requirements practitioners need to meet.

FrameworkControl / ReferenceRelevance
OWASP Non-Human Identity Top 10NHI-03 — Vulnerable Third-Party NHIRepository-supplied filter logic acted like an untrusted third-party control plane for the assistant.
Recommendation — Review any repository-originated configuration that can change AI-visible output under NHI-03.
OWASP Agentic AI Top 10ASI03 — Identity & Privilege AbuseThe attack abused the assistant's trusted visibility rather than its execution path.
Recommendation — Map AI pipeline trust boundaries to ASI03 and restrict who can alter model-visible context.
MITRE ATT&CKTA0003;TA0006;TA0007 — Persistence; Credential Access; DiscoveryThe technique hid malicious code and suppressed review evidence across the workflow.
Recommendation — Map hidden-review behaviour to TA0003, TA0006, and TA0007 to improve detection coverage.
NIST CSF 2.0PR.AA-05 — Access Permissions, Entitlements and AuthorizationsThe tool granted repository-supplied settings authority without a separate trust decision.
Recommendation — Apply PR.AA-05 to ensure AI pipeline permissions distinguish user-controlled from repository-controlled inputs.
OWASP API Security Top 10API8 — Security MisconfigurationAutomatic loading of untrusted project configuration is a misconfiguration of trust handling.
Recommendation — Harden configuration handling under API8-style misconfiguration controls for tool inputs.

Key terms

  • Observation Layer: The observation layer is the part of an AI workflow that determines what information reaches the model for reasoning and review. It includes preprocessors, filters, hooks, and connectors. If this layer is compromised, the model may appear accurate while operating on a deliberately incomplete view.
  • Trust Laundering: Trust laundering is when untrusted content gains trusted authority simply by passing through a tool that assumes the source is safe. In AI-assisted development, that can happen when repository files or hooks silently shape what the model sees, turning evidence selection into a security control.
  • Perceptual Authority: Perceptual authority is the power to decide what an AI or reviewer is allowed to see before making a judgement. Unlike execution authority, it does not run code, but it can still shape decisions by suppressing warnings, diffs, or malicious lines from the evidence stream.
  • Project-Local Configuration: Project-local configuration is settings stored in a repository and loaded automatically by tooling during work on that codebase. It is useful for shared behaviour, but it becomes risky when loaded without provenance checks, because attacker-controlled settings can inherit authority they were never meant to have.

Deepen your knowledge

NHI governance, agentic AI identity, and machine identity lifecycle are core topics in our NHI Foundation Level course, the industry's only accredited NHI security programme. If you are building or maturing an IAM programme, it is worth exploring.
NHIMG Editorial Note
Published by the NHIMG editorial team on June 9, 2026.
Updated on October 10, 2026.
NHI Mgmt Group, the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org