Join our Newsletter — 33% off our NHI Course

Context-Format-Salience Model

A way to explain why some indirect prompt injections work and others fail. Context is whether the payload matches the task, format is whether it blends into the medium, and salience is whether it captures the model’s attention strongly enough to influence action.

What the model is actually explaining

The Context-Format-Salience Model is a prompt-injection lens, not a general AI theory. It says an indirect injection is more likely to work when the malicious content fits the surrounding task, looks native to the channel, and is prominent enough to pull the model away from the intended objective.

That makes it useful for thinking about why one hidden instruction is ignored while another is followed. The model does not require the payload to be overtly malicious, only to be compatible with the surrounding context and difficult for the system to mentally discount.

Context: why task fit matters

Context is the degree to which the injected instruction aligns with what the model is already doing. If a payload looks like a natural continuation of the user’s task, or borrows the same vocabulary and goal structure, it can be easier for the model to treat it as relevant rather than foreign.

This is why prompt injection often succeeds through framing instead of force. The payload does not need to override the task directly, it only needs to appear plausibly related enough that the model blends it into its active reasoning. In agentic workflows, that can be especially dangerous when the model is chaining steps across tools or message sources, as described in NHIMG’s OWASP Agentic Applications Top 10.

Format: why the medium shapes susceptibility

Format refers to whether the payload blends into the delivery channel. An instruction hidden in plain text, HTML-like structure, metadata, comments, tables, or other structured content can appear less like an attack and more like ordinary input.

The model is therefore not just reading words, it is interpreting presentation. A payload that matches the native style of the medium can survive weak filtering, evade human review, and enter the model’s attention as if it were part of the trusted material. That is why protocol and interface boundaries matter, including cases where tool-facing agents interact with structured messages or remote services, a topic covered in the MCP Security Guide.

Salience: why attention can be redirected

Salience is the strength of the payload’s pull on the model’s attention. Highly salient text can be emotionally loaded, directive, repetitive, urgent, or otherwise attention-grabbing, making it more likely to influence the model than quieter surrounding instructions.

This part of the model helps explain why some injections do not need to be sophisticated. They just need to be hard to ignore. In practice, salience can compete with the model’s intended task priority, especially when the system lacks strong instruction hierarchy, source separation, or robust trust boundaries.

Why the three factors work together

The model is strongest when context, format, and salience reinforce one another. A payload that is task-aligned, visually native to the channel, and highly attention-grabbing has a better chance of being treated as legitimate input rather than adversarial interference.

That combination is what makes indirect prompt injection so hard to eliminate with simple keyword filters. Security has to account for where the instruction appears, how it is packaged, and how strongly it can compete for model attention, not just whether it contains obviously malicious phrasing.

Risk and Threat Considerations

This model highlights a real trust problem for AI systems that consume untrusted text, documents, web pages, tickets, emails, or tool outputs. When context, format, and salience line up, indirect prompt injection can cause instruction hijacking, policy bypass, or unsafe tool use without the attacker ever touching the core prompt.

Failure mechanism: The model mistakes injected content for relevant task material because the payload fits the surrounding context, blends into the format, and captures attention more strongly than the intended instruction hierarchy.

Impact: The system may disclose sensitive information, alter outputs, call tools incorrectly, or carry out attacker-directed actions inside an agentic workflow.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

OWASP Agentic AI Top 10 addresses the attack and risk surface, while NIST AI RMF sets the governance and control requirements practitioners need to meet.

Framework Control / Reference Relevance
OWASP Agentic AI Top 10 ASI01 — Agent Goal Hijack Indirect prompt injection redirects an agent from its intended goal.
ASI02 — Tool Misuse Injected text can induce unsafe tool calls or tool-using behaviour.
ASI06 — Memory & Context Poisoning The model’s context can be polluted by hostile input that alters behaviour.
Recommendation — Separate untrusted content from instructions and constrain goal-changing inputs. Restrict tool invocation to approved intents and validate every tool action. Filter and isolate external context before it reaches decision-making memory.
NIST AI RMF GOVERN — AI Risk Management Governance The model is a governance lens for managing AI injection risk.
Recommendation — Document AI input trust boundaries and assign ownership for injection risk.

Practitioner Guidance

What to watch for: Treat this model as a design warning, not a detection rule. The practical signal is any workflow that ingests external text and then lets the model act on it without clear separation between untrusted content and executable instruction.

Practitioner takeaway: Stronger prompt hygiene helps, but the real control objective is to reduce the model’s ability to confuse content with instruction and to limit what a diverted instruction can do.