Runtime filters can reduce harmful prompts and outputs, but they do not solve upstream poisoning, dependency compromise, or over-scoped tool access. If the model can still reach privileged systems or inherit unsafe context, the control is only covering the last step of the chain, not the conditions that created the risk.
What runtime filters do protect, and what they leave exposed
Runtime filters are useful as a final safety layer: they can block obvious harmful prompts, reduce toxic or disallowed outputs, and catch some policy violations before the user sees them. The break point is that they operate at the end of the chain. If the model has already been steered by poisoned context, compromised dependencies, or excessive tool permissions, the filter sees only the symptom, not the cause.
That matters because many LLM failures are upstream of the generated text itself. A safe-looking response can still be built on tainted retrieval, injected instructions, or a malicious plugin path. In those cases, the filter may suppress the final answer, but it does not restore trust in the context, the tool call, or the system state that produced it.
Runtime filtering is therefore best understood as containment, not prevention. It can reduce exposure from unsafe output, but it does not prove the model is operating on clean inputs or bounded authority. If the surrounding system is over-permissive, the control is reactive rather than preventative.
Where the real failure happens: context, dependencies, and tools
The strongest single weakness of a filter-only design is that it treats the model output as the main risk surface. In practice, the more important surfaces are the prompts, retrieval sources, model supply chain, and connected tools. If any of those are compromised, the model may be prompted into unsafe behavior before the runtime filter ever evaluates the result.
That is why upstream poisoning is so damaging. Poisoned retrieval can bias answers, malicious memory can persist across sessions, and compromised dependencies can alter model behaviour without changing the final filter logic. A filter cannot reliably distinguish benign output from output that was shaped by a corrupted trust path.
Tool access is the other major blind spot. If an LLM can call email, file systems, ticketing platforms, cloud APIs, or admin consoles, then authorization becomes part of the security boundary. A runtime filter does not stop the model from being handed more privilege than it should have in the first place.
What has to be controlled before the filter ever runs
The control stack has to start with trust boundaries, not moderation. Clean retrieval, verified dependencies, scoped tool permissions, and explicit approval gates matter more than downstream output filtering when the model can trigger actions or access sensitive data. A filter can support this stack, but it cannot substitute for it.
That is especially true in systems that blend LLMs with search, memory, and automation. Once the model can read from or act on privileged systems, the security question changes from “Did the output look safe?” to “Was the model ever allowed to see, infer, or execute something unsafe?” Runtime filtering only answers the first question.
For practitioners, the operational test is whether the model can still cause harm even when every visible output is screened. If the answer is yes, the real fix is to reduce what the model can reach, what it can remember, and what it can do by default.
Risk and Threat Considerations
Runtime filters can create a false sense of safety when the underlying system still trusts unverified inputs or grants broad access. The risk is not just harmful output, it is hidden compromise: poisoned context, dependency tampering, or over-scoped tools can turn the model into a conduit for data exposure or unauthorized action.
Failure mechanism: An attacker or faulty integration alters the prompt chain, retrieval layer, package dependency, or tool path, so the model produces or triggers unsafe behavior before the filter has any meaningful chance to correct the underlying state.
Impact: The organisation may block some bad outputs while still exposing sensitive data, executing unwanted actions, or inheriting compromised decisions from a tainted upstream control plane.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Agentic AI Top 10, OWASP Non-Human Identity Top 10 and MITRE ATT&CK address the attack and risk surface, while NIST SP 800-53 Rev 5 and NIST AI RMF set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP Agentic AI Top 10 | ASI03 — Identity & Privilege Abuse | Runtime filters cannot stop over-scoped agent/tool authority. |
| Recommendation — Constrain tool and privilege scope before relying on output filtering. | ||
| OWASP Non-Human Identity Top 10 | NHI-02 — Secret Leakage | Upstream compromise can expose secrets before any runtime filter sees output. |
| NHI-05 — Overprivileged NHI | Over-scoped non-human identities let LLMs do damage beyond filtered output. | |
| NHI-09 — NHI Reuse | Shared identities and credentials increase the blast radius of upstream compromise. | |
| Recommendation — Protect secrets in prompts, retrieval, and tool-connected systems. Reduce machine and service privileges before deploying runtime filters. Eliminate reused identities and credentials across AI-connected services. | ||
| NIST SP 800-53 Rev 5 | AC-6 — Least Privilege | The question centers on excessive model access beyond output moderation. |
| SI-4 — System Monitoring | Runtime filters are detection-like; upstream compromise needs broader monitoring. | |
| IA-9 — Identification and Authentication (Non-Organizational Users) | Model-to-service access and delegated tool use depend on controlled authentication. | |
| Recommendation — Limit LLM-connected accounts and tools to the minimum needed access. Monitor upstream inputs, tool calls, and dependency behavior for abuse. Use strong service authentication for every tool and dependency connection. | ||
| NIST AI RMF | GV.4 — AI governance | The issue is governance of end-to-end AI risk, not just output moderation. |
| Recommendation — Govern the full AI lifecycle, including inputs, tools, and deployment boundaries. | ||
| MITRE ATT&CK | T1552 — Unsecured Credentials | The subject includes stolen or exposed credentials as an upstream failure path. |
| T1078 — Valid Accounts | Over-scoped access turns compromised accounts into a direct abuse path. | |
| Recommendation — Hunt for credential exposure in logs, prompts, and connected systems. Detect and limit abuse of valid accounts used by AI-connected tooling. | ||
Practitioner Guidance
What to prioritise: Treat runtime filters as one control in a larger trust architecture. The higher-value work is to constrain tool scope, isolate retrieval and memory, and verify upstream inputs before the model can act on them.
What to verify: Check whether the model can reach privileged systems, whether those connections are explicitly approved, and whether a blocked output would still leave behind a dangerous tool invocation, data read, or state change.
Common mistake: Teams often evaluate only prompt safety and moderation effectiveness. That misses the more important question of whether the LLM is authorised to touch systems whose compromise would matter more than any single unsafe response.
Practitioner takeaway: If the model still has unsafe reach, a runtime filter is a guardrail, not a boundary, and the boundary must be moved upstream to data, dependency, and access control.