Keyword-only filters fail when an attacker asks indirectly, translates the request, or splits the secret into fragments the model can still reconstruct. The guard sees safe text, but the model still understands the disclosure intent. That is why prompt injection often succeeds even when obvious secret terms are blocked.
Why keyword filters fail against prompt injection
Keyword-only filtering is brittle because it treats surface text as the security boundary instead of the model’s interpreted intent. Attackers can disguise harmful requests by using indirect phrasing, translation, encoding, or fragmenting sensitive terms across multiple turns. The filter sees harmless tokens, but the model still reconstructs the forbidden meaning.
That matters because prompt injection is often an interpretation problem, not a vocabulary problem. A guardrail that blocks obvious secret names can still miss paraphrases, oblique requests, and multi-step prompts that achieve the same disclosure outcome without ever matching the banned word list.
When the control depends on exact terms, it also creates a false sense of safety. Teams may believe they have blocked “secret” or “credential” leakage, but the model can still be steered into revealing the same information through context, transformation, or decomposition.
How attackers bypass text-only prompt filters
Bypass usually works by separating the dangerous intent from the visible trigger. Common patterns include asking for a translation, requesting a summary of benign-looking fragments, or distributing the sensitive request across multiple messages so the model assembles the meaning only after the filter has already passed each piece.
This is why the relevant security problem is not just keyword matching, but instruction hierarchy and contextual trust. The system must decide whether a prompt is allowed based on its effective meaning, not on whether one prohibited term appears in the raw string.
Text-only filters are also weak against adversarial paraphrase. As with classic content moderation failures, the attacker only needs one alternate path that preserves the objective while evading the literal denylist. If the control is not semantically aware, it will be bypassed eventually.
What a defensible prompt-defense layer has to check instead
A practical control stack needs layered checks: semantic classification, policy enforcement, context isolation, and output monitoring. Keyword lists can still help as a cheap first pass, but they should be treated as a signal, not as the decision rule.
For AI applications that accept external text, NIST AI 600-1 GenAI Profile is useful because it frames prompt safety as part of broader GenAI risk management rather than a simple denylist problem. The model should be evaluated for instruction-following robustness, disclosure resistance, and safe handling of untrusted input.
Where the system includes tool use, memory, or delegated actions, the filter must also account for what the model can do after the prompt is accepted. A request that looks harmless at the text layer can still become dangerous if it can trigger retrieval, tool calls, or data export downstream. OWASP Agentic AI Top 10 and CSA MAESTRO agentic AI threat modeling framework both help teams reason about that broader execution path.
Risk and Threat Considerations
Keyword-only filters create a disclosure gap: the attacker does not need prohibited words if the model can infer the intent from context, fragments, or transformed text. That makes secret extraction, policy bypass, and harmful instruction injection more likely than teams expect.
Failure mechanism: The control validates literal tokens instead of the model’s interpreted request, so paraphrase, translation, multi-turn decomposition, and encoded input can evade the filter while preserving the malicious objective.
Impact: Sensitive data can be revealed, unsafe actions can be triggered, and operators can wrongly trust a control that has not actually reduced the model’s effective attack surface.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Agentic AI Top 10 and CSA MAESTRO address the attack and risk surface, while NIST AI RMF and OWASP ASVS set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| NIST AI RMF | Generative AI risk management | Prompt filtering is part of GenAI risk control and adversarial robustness. |
| Recommendation — Evaluate prompts for semantic bypasses and manage disclosure risk across the full GenAI lifecycle. | ||
| OWASP Agentic AI Top 10 | ASI06 — Memory & Context Poisoning | Indirect prompt attacks exploit context and instruction handling, not just keywords. |
| ASI01 — Agent Goal Hijack | Attackers steer systems toward unsafe goals through disguised instructions. | |
| Recommendation — Harden context handling and test for prompt injection that survives literal filtering. Validate that policy checks block hostile objectives, not just banned terms. | ||
| CSA MAESTRO | Agentic AI threat modeling | Threat modeling helps assess prompt abuse, tool use, and downstream abuse paths. |
| Recommendation — Model prompt abuse paths end to end, including downstream tool and data exposure. | ||
| OWASP ASVS | V1 — Encoding and Sanitization | Untrusted input handling must account for transformed or obfuscated text. |
| Recommendation — Normalize and inspect untrusted input before relying on content filters. | ||
Practitioner Guidance
What to prioritise: Treat prompt filtering as one layer inside a broader input and output governance design. The first priority is to stop assuming that denylisted words are equivalent to denied intent.
What to verify: Test the control with paraphrases, translation, fragmented prompts, and multi-turn attacks. If the model still reaches the same unsafe answer path, the filter is not enforcing the policy you think it is.
Common mistake: Teams often overinvest in ever-larger keyword lists and underinvest in semantic evaluation, context separation, and post-generation checks. That usually creates brittle coverage without materially improving safety.
Practitioner takeaway: If the model can understand the request, a keyword filter is already too shallow to trust on its own.
Related resources from NHI Mgmt Group
Deepen Your Knowledge
Free weekly newsletter
Subscribe to the NHI & AI Identity Journal
The latest on NHI and Agentic AI security – articles, research, breaches, news and events every week.
Bonus 33% off our NHI Course when you subscribe.
Reviewed and updated by the NHIMG editorial team on October 11, 2026.
NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org