Join our Newsletter — 33% off our NHI Course

How should security teams test guardrails against prompt injection that uses formatting instead of direct disclosure?

Test for reconstructed intent, not just forbidden wording. A guardrail can block obvious jailbreak phrases and still leak a secret when the same request is reframed as formatting, encoding, translation, or per-character output. Defenders should normalize the model output, inspect whether it reconstructs protected data, and treat transformed disclosures as a policy failure.

How to test for formatting-based prompt injection

Guardrails should be tested against what the model ultimately reconstructs, not only the exact forbidden phrasing the attacker used. A request can hide the same disclosure inside spacing, code blocks, character-by-character output, transliteration, or other formatting tricks. The right test is whether the output can be normalized into protected content, not whether it looks obviously malicious at first glance.

That means a strong test case set should include paraphrase, encoding, fragmentation, and reconstruction attempts. If the guardrail only stops direct requests for a secret but allows the same secret to be reassembled through formatting, the control is too narrow.

What a good evaluation harness should check

A useful harness normalizes candidate output before judging it. Teams should compare the normalized result against protected strings, sensitive instructions, and policy-bound content, then flag cases where formatting changes preserve the underlying disclosure. This is especially important for assistant workflows that produce markdown, tables, lists, transcripts, or per-character renderings.

It also helps to test both input and output pathways. Some attacks try to smuggle intent into the prompt through layout, while others try to get the model to emit a harmless-looking transformation that becomes sensitive once reassembled. A solid harness exercises both directions so the guardrail is not only checking the front door.

For agentic systems, this is the difference between blocking a bad sentence and blocking a bad action path. Agentic AI Security Guide is useful here because it frames prompt injection alongside tool misuse, memory abuse, and other runtime control failures that are often tested through indirect prompts rather than obvious jailbreaks. For red-team style validation, Red Teaming AI Agents for Identity Abuse helps teams structure tests around privilege misuse, delegation abuse, and exfiltration paths that a formatting-only payload may still trigger.

Why formatting-based attacks slip past weak guardrails

Many guardrails are built as keyword filters or pattern matchers, so they overfit to obvious jailbreak language. An attacker can ask for the same information as a translation, a cipher, a markdown artifact, a one-character-per-line response, or a cleanup task. If the model complies, the policy failure sits in the reconstruction step, not in the surface text.

This is why test cases should include intent-preserving transformations. A prompt injection is still successful if the answer can be decoded, concatenated, or normalized into a protected instruction or secret. Strong defenses need to judge the semantic end state of the output, not just the literal tokens used along the way.

Risk and Threat Considerations

Formatting-based prompt injection creates a false sense of safety because the response can look compliant while still leaking sensitive material after trivial normalization. In practice, that makes it harder for humans, log review, and simple string-based detectors to spot the failure.

Failure mechanism: the attacker shifts the request into a transformed output path, such as encoding, translation, segmentation, or structured formatting, so the guardrail evaluates harmless-looking text while the underlying content still reconstructs to protected data.

Impact: sensitive instructions, secrets, or policy-restricted content can be disclosed without triggering naive filters, which weakens red-team coverage, undermines incident review, and increases the chance that a malicious prompt survives into production unnoticed.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

OWASP Agentic AI Top 10, MITRE ATT&CK and OWASP Non-Human Identity Top 10 address the attack and risk surface, while NIST AI RMF sets the governance and control requirements practitioners need to meet.

Framework Control / Reference Relevance
OWASP Agentic AI Top 10 ASI01 — Agent Goal Hijack Formatting-based prompt injection still hijacks the agent's output goal.
ASI02 — Tool Misuse The attack can force unsafe transformations through agent actions or tools.
ASI09 — Human-Agent Trust Exploitation The guardrail can be bypassed by making unsafe output look innocuous to reviewers.
Recommendation — Test prompts that try to redirect output into hidden or reconstructed disclosures. Validate that tools cannot be used to transform benign text into a policy breach. Review outputs for disguised policy violations that would mislead a human reviewer.
MITRE ATT&CK Adversarial Tactics, Techniques, and Procedures Prompt injection with encoding or formatting is a practical adversary technique.
Recommendation — Map transformed-disclosure tests to observed adversary prompt-injection techniques.
NIST AI RMF AI Risk Management Framework The topic is AI guardrail testing and semantic validation of model outputs.
Recommendation — Evaluate AI outputs for harmful reconstruction after normalization and transformation.
OWASP Non-Human Identity Top 10 NHI-02 — Secret Leakage The question concerns leaking protected data through transformed output.
NHI-10 — Human Use of NHI Guardrail testing must account for humans reviewing or reusing transformed outputs.
Recommendation — Add tests that detect secret leakage even when the secret is formatted or encoded. Block workflows that let humans reintroduce transformed sensitive output into production.

Practitioner Guidance

What to verify: Build assertions around normalized output, not raw text. If your evaluator cannot reliably reconstruct whether the model has leaked protected content after formatting changes, the test is incomplete.

Common mistake: Treating successful block rates on obvious jailbreak language as proof of safety. A guardrail is not mature until it also fails closed on transformed disclosures that preserve the same intent.

What good looks like: The system rejects or sanitizes requests that would reconstruct sensitive content, and your test suite demonstrates that encoding, spacing, per-character output, and translation do not bypass the policy decision.

Practitioner takeaway: Measure the guardrail against semantic recovery, not prompt appearance, because attackers will often hide the violation in the transformation rather than the wording.