Join our Newsletter — 33% off our NHI Course

What do teams get wrong when they assume AI models can safely handle untrusted text, emojis, or encoded characters?

Teams often assume that obvious content filters are enough, but adversaries can hide instructions in terminal-like framing, Unicode smuggling, emojis, or other contextual tricks. These inputs can bypass simplistic checks and change how the model interprets intent. A stronger approach is layered validation, normalisation, and testing against multiple encoding and framing techniques before deployment.

Why simplistic filters fail on hidden instructions

The mistake is treating untrusted text as if the visible surface is the whole threat. AI systems can be steered by instructions that are disguised with punctuation, terminal-style framing, Unicode confusables, zero-width characters, emoji sequences, or mixed encodings. If the model or preprocessor normalises those inputs inconsistently, the attack can survive the filter and still influence interpretation.

That is why content safety has to be handled as a parsing and trust-boundary problem, not just a keyword problem. Normalisation, canonicalisation, and consistent decoding need to happen before policy checks, and the checks need to be designed around the exact forms the model will actually see. For broader agent and model threat patterns, MITRE ATLAS adversarial AI threat matrix is a useful reference, because it captures prompt injection, context manipulation, and related abuse patterns.

Teams also underestimate how often the failure is not a single bypass but a chain: one layer strips or rewrites characters, another layer preserves them, and the model receives a version that no reviewer intended. That is why layered validation matters more than one “safe input” rule.

What goes wrong with Unicode, emojis, and encoded characters

Unicode and encoding issues create ambiguity that attackers can exploit. The same visual string may contain different code points, hidden direction changes, or text that renders one way to humans and another way to the model. Emojis and special symbols can also act as framing devices, separators, or attention anchors that change the model’s reading of nearby content.

Encoded payloads create a second class of failure. Base64, percent encoding, HTML entities, and nested escapes can hide instruction fragments until after a decoder, proxy, or middleware expands them. If each layer decodes differently, the security control may inspect one representation while the model consumes another. This is especially dangerous when the system mixes sanitisation for display with sanitisation for inference.

A practical example is trusting that “non-textual” characters are inert. They are not. They can alter tokenisation, break pattern matching, or create context boundaries that the model treats as meaningful. NIST AI Risk Management Framework is relevant here because it frames this as a trustworthy AI risk issue, where input handling, robustness, and monitoring all affect system behaviour.

How teams should test and harden the input path

The right control mindset is to test the full input path, not just the prompt template. Inputs should be normalised to a canonical form, rejected when they contain unsafe or unexpected control characters, and re-tested after every transformation step so that preprocessing does not reintroduce risk. When the model is part of a larger workflow, the safe boundary has to include the wrapper, middleware, and any retrieval or tool layer that reuses the text.

Security testing should include adversarial cases that look benign to a human reviewer but behave differently after decoding or rendering. That means checking for mixed encodings, Unicode edge cases, zero-width characters, bidi controls, emoji obfuscation, and nested framing. The goal is not to ban all rich text, it is to prove that dangerous transformations cannot change intent or execution.

For teams building governed AI programmes, ISO/IEC 42001:2023 AI Management System Standard helps anchor this work in repeatable controls around accountability, review, and continual improvement. If the system accepts user content that can influence downstream actions, the validation logic should be treated as a control surface, not a convenience feature.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

NIST AI RMF sets the technical controls, while ISO/IEC 42001:2023 defines the regulatory obligations.

Framework Control / Reference Relevance
NIST AI RMF Govern Input normalization and robustness are core to trustworthy AI risk management.
Recommendation — Establish input-handling controls that preserve trust boundaries across preprocessing and inference.
ISO/IEC 42001:2023 AI management system requirements The question is about governing AI system behaviour under adversarial inputs.
Recommendation — Define review, accountability, and continual-improvement controls for unsafe input handling.

Practitioner Guidance

What to prioritise: Validate and normalise before the model ever sees the text, then recheck the exact post-processed form that will be consumed by the model or agent. If the security decision depends on string matching alone, assume it will be bypassed.

What to verify: Confirm that the same input produces the same canonical form across the web tier, API layer, logging pipeline, and model wrapper. A mismatch between what reviewers see and what the model receives is a sign that the control path is not trustworthy.

Common mistake: Treating emojis, Unicode, and encoded text as presentation issues instead of security issues. In practice, they can change tokenisation, context, and downstream behaviour, so they need the same scrutiny as any other untrusted input.

Practitioner takeaway: If untrusted text can influence model behaviour, the real control is not “blocking bad words”, it is proving that every representation, decode step, and transform preserves the intended security boundary.