They can miss payloads that change character rendering or token boundaries without changing the attacker’s intended meaning. In practice, the classifier and the model may process different representations of the same input, which lets malicious instructions slip through a guardrail even when they look obvious to a human reviewer.
Why raw-text-only inspection fails
Prompt-injection filters that only read raw text assume the same bytes, the same visible text, and the same model input are always equivalent. They are not. Attackers can exploit rendering layers, encoding transforms, normalization rules, or tokenization differences so the filter sees one thing while the classifier or model effectively processes another.
The practical break is representation mismatch. A guardrail built on literal text comparison can miss instructions hidden by Unicode shaping, zero-width characters, markup, alternate encodings, or other transformations that preserve attacker intent while changing the surface form. That makes “looks safe in the filter” a poor indicator of what the downstream model will actually interpret.
Once a system compares only the raw string, it also loses visibility into where the trust boundary really sits. The security decision is then made before the content has been canonicalized, tokenized, or rendered in the same way the model will consume it, which is exactly where prompt-injection payloads can survive.
What the attacker is exploiting
Prompt injection succeeds when the defender validates one representation and executes another. The attacker is not relying on obvious keyword evasion alone; they are relying on differences in normalization, parsing, or display so the payload remains semantically intact while escaping simple inspection.
This is why indirect prompt injection is especially hard to catch with raw-text rules. A payload can be embedded in content that appears benign to a reviewer, yet still influence the model once the surrounding system strips, rewrites, or reinterprets the input. In agentic workflows, that gap can reach tool use, memory, or delegated actions, not just chat output.
For a concrete example of how hidden instructions and representation tricks can be abused in agent workflows, see Agentic AI Security Guide and Browser and Computer-Use Agent Security Guide. Both show why the inspection point has to match the model’s actual consumption path, not just the visible text stream.
How defenders should think about the control gap
The control gap is not “missing a bad phrase,” it is failing to inspect the same canonical form that the model will process. If a filter, renderer, tokenizer, or downstream agent sees content differently, the guardrail becomes bypassable by design rather than by luck.
That matters most where content can drive actions. When hidden or altered instructions reach a tool-enabled model, the effect is no longer just unsafe wording, it can become unauthorized retrieval, data disclosure, or unwanted execution. Stronger defenses therefore focus on normalization, structured parsing, context separation, and policy enforcement after transformation, not before it.
Representative attack cases make the point: EchoLeak (Microsoft 365 Copilot) 2025 shows zero-click prompt injection against assistant context, while Sentry MCP Agentjacking 2026 shows how deceptive tool output can push coding agents into executing attacker-controlled actions.
Risk and Threat Considerations
Raw-text-only filtering creates a false sense of safety because the most dangerous prompt-injection payloads often survive by changing form, not meaning. Once the defender and the model are no longer evaluating the same normalized content, the attacker can route instructions around the guardrail while keeping the payload understandable to the downstream system.
Failure mechanism: The filter inspects a surface string, but the model consumes a transformed representation after normalization, rendering, tokenization, or parser changes. That mismatch lets adversarial instructions bypass detection even when a human reviewer can spot them in the rendered view.
Impact: Malicious instructions can slip into retrieval, tool use, memory, or agent actions, leading to data exfiltration, unauthorized operations, or broader compromise of trusted workflows.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Agentic AI Top 10 and MITRE ATT&CK address the attack and risk surface, while NIST AI RMF sets the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP Agentic AI Top 10 | ASI03 — Identity & Privilege Abuse | Raw-text bypasses can drive unauthorized agent actions through trusted identity and privilege channels. |
| ASI02 — Tool Misuse | Prompt injections often aim to coerce unsafe tool calls after deceptive input transformation. | |
| ASI06 — Memory & Context Poisoning | Hidden or transformed payloads can poison the context a model later relies on. | |
| Recommendation — Constrain agent privileges and verify input provenance before allowing actions. Validate tool-intent boundaries and block untrusted instructions from tool execution. Sanitize and compartmentalize context before it is stored or reused. | ||
| MITRE ATT&CK | T1056 — Input Capture | Prompt injection relies on tampering with the content captured and interpreted by a target system. |
| Recommendation — Inspect the full input path for transformations that alter security decisions. | ||
| NIST AI RMF | Map, Measure, and Manage AI Risks | The issue is an AI risk-management problem involving model/input mismatch and guardrail failure. |
| Recommendation — Measure whether model-facing content matches the content your controls actually inspected. | ||
Practitioner Guidance
What to verify: Check whether your guardrail evaluates the same canonical input form that the model and any downstream tools actually use. If the answer is no, treat the control as incomplete even if it catches obvious direct-text payloads.
Decision rule: If a prompt can be transformed by Unicode, markup, encoding, or rendering changes, validate after canonicalization and before action, not just at ingestion. If the system cannot guarantee that equivalence, assume the filter is bypassable.
Practitioner takeaway: Prompt-injection defense fails when inspection and execution happen on different representations, so the control objective is representation equivalence, not keyword detection.
Related resources from NHI Mgmt Group
- What breaks when prompt injection guardrails only look for obvious malicious text?
- Why do prompt injection defenses fail when they only inspect the text?
- What breaks when prompt injection controls only inspect user prompts and not retrieved content?
- What breaks when a browser guardrail only sees the surface text of prompt injection?
Deepen Your Knowledge
Free weekly newsletter
Subscribe to the NHI & AI Identity Journal
The latest on NHI and Agentic AI security – articles, research, breaches, news and events every week.
Bonus 33% off our NHI Course when you subscribe.
Reviewed and updated by the NHIMG editorial team on October 11, 2026.
NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org