A prompt attack that uses examples inside the context window to steer a model into producing harmful outputs under a learned task pattern. It relies on the model completing the structure of the examples, rather than being deceived by encoding, role play, or hidden instructions.
What Involuntary In-Context Learning Is
Involuntary in-context learning is a prompt attack that exploits the model’s tendency to infer and continue a pattern from examples already present in the context window. The attacker is not trying to hide instructions, but to make the model complete the learned structure in a harmful direction.
How the Attack Works
The attack depends on the model treating nearby examples as a task template. Once the prompt establishes a pattern, the model may generalise from that pattern instead of applying the user’s intended instruction, especially when the examples are repetitive, strongly formatted, or semantically suggestive.
This makes the attack different from classic prompt injection that relies on direct instruction override. Here, the harmful behaviour is induced through completion pressure: the model imitates what it thinks the task requires because the context itself implies the shape of the answer.
That pattern-following behaviour is why the issue matters in prompt-heavy systems, evaluation pipelines, and chained LLM workflows. A seemingly benign example set can steer the model into producing unsafe content, policy leakage, distorted summaries, or outputs that reinforce the attacker’s framing.
Where It Shows Up in Real Systems
Involuntary in-context learning is most likely where a model receives long prompts, retrieved examples, embedded transcripts, or multi-turn context that mixes instructions with sample outputs. It can also appear when a system reuses prior context across sessions or composes prompts from untrusted sources.
In practice, the risk grows when the application assumes that examples are neutral reference material. A model may still absorb those examples as behavioural guidance, especially if they are placed close to the desired output format or repeated in a way that makes the malicious pattern feel authoritative.
The security challenge is not limited to textual harm. The same mechanism can be used to bias classification, leak hidden policies, distort compliance responses, or make the model mirror an attacker’s structure in a way that changes downstream automation.
How to Recognise and Defend Against It
Defence starts with treating examples as untrusted input, not just content. Strong prompt boundaries, clear role separation, and minimising exposure to attacker-controlled examples all reduce the chance that the model will infer the wrong task from the surrounding context.
It also helps to test systems for pattern sensitivity, because the failure mode is often subtle. Models may appear to follow the instruction correctly until the context is shaped to encourage imitation, at which point the output quality or safety degrades in a way that is easy to miss without adversarial testing.
For linked references on the broader attack family, see MITRE ATLAS adversarial AI threat matrix, which maps context poisoning and related AI abuse techniques, and OWASP Agentic AI Top 10, which covers memory and context poisoning risks in agentic systems.
Why It Matters for Prompt Security
Involuntary in-context learning shows that prompt security is not only about stopping explicit malicious instructions. It is also about controlling the latent behaviour the model extracts from examples, formatting, and repetition inside the context window.
That makes prompt design a security control, not just a usability concern. Systems that rely on retrieval, demonstrations, or generated examples need to assume that the model can be nudged by the very material intended to help it perform better.
For practitioners building defence-in-depth around model interaction, the issue is closely related to broader context integrity and model steering risk. See also NIST AI Risk Management Framework for governance of AI risk, and NIST SP 800-53 Rev 5 Security and Privacy Controls for control families that support access control, integrity, and monitoring.
Risk and Threat Considerations
Involuntary in-context learning can turn normal examples into an attack surface. If an attacker can shape the context, they may steer the model toward unsafe, policy-violating, or misleading outputs without needing to overwrite the user’s instruction directly.
Failure mechanism: The model generalises from the examples in the prompt and completes the pattern it believes is being demonstrated, so malicious structure can dominate the intended task.
Impact: The result can be harmful content generation, incorrect automated decisions, hidden policy leakage, or a prompt-mediated compromise of output integrity.
MITRE ATLAS adversarial AI threat matrix helps place this failure mode alongside context poisoning and other adversarial AI techniques.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
MITRE ATT&CK, OWASP Agentic AI Top 10 and OWASP API Security Top 10 address the attack and risk surface, while NIST AI RMF and NIST SP 800-53 Rev 5 set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| MITRE ATT&CK | T1059 — Command and Scripting Interpreter | Captures adversary-driven prompt and context abuse as a technique family |
| T1565 — Data Manipulation | Relevant where prompt context is altered to distort model outputs | |
| Recommendation — Map prompt-pattern abuse to T1059-adjacent tradecraft and monitor for malicious instruction shaping. Detect and block context tampering that changes downstream model behaviour. | ||
| NIST AI RMF | GOVERN — Govern | Applies to managing AI risk, accountability, and controls around prompt-driven systems |
| MAP — Map | Supports identifying context-window abuse as an AI risk scenario | |
| MEASURE — Measure | Supports testing how prompts and examples steer model behaviour | |
| Recommendation — Establish governance for prompt integrity, adversarial testing, and model-use boundaries. Document where context poisoning and example-steering can affect model outputs. Measure prompt sensitivity with adversarial examples and red-team exercises. | ||
| OWASP Agentic AI Top 10 | ASI06 — Memory & Context Poisoning | Directly covers harmful context shaping that alters model or agent behaviour |
| ASI09 — Human-Agent Trust Exploitation | Relevant where users or systems trust examples that secretly steer the agent | |
| ASI10 — Rogue Agents | Supports defending against autonomous or semi-autonomous misuse of model behaviour | |
| Recommendation — Isolate untrusted context and test for context-poisoning-driven output drift. Restrict trust in retrieved examples and validate whether context is attacker-controlled. Constrain autonomous action paths when model outputs can be manipulated by context. | ||
| OWASP API Security Top 10 | API6 — Unrestricted Access to Sensitive Business Flows | Applies when prompt steering causes unintended access to sensitive workflows |
| Recommendation — Gate sensitive flows so model outputs cannot trigger privileged business actions. | ||
| NIST SP 800-53 Rev 5 | SI-10 — Information Input Validation | Directly supports validating untrusted prompt and context input before use |
| Recommendation — Validate and sanitise context inputs before they can influence model behaviour. | ||