Context-jacking is the abuse of an AI model’s conversation or retrieval context to redirect its behaviour away from intended policy. Attackers exploit hidden instructions, embedded data, or manipulated memory to influence outputs and actions. The risk rises when systems trust context without validating source, intent, and scope.
Expanded Definition
Context-jacking describes a class of prompt and retrieval abuse where an AI system is influenced by surrounding context rather than by the user’s intended instruction. That context may include hidden prompts, retrieved documents, long-conversation history, tool outputs, or memory objects that the system treats as authoritative.
The boundary matters: context-jacking is not simply “bad prompting” and not every model error is an attack. The term is used when an actor deliberately manipulates the contextual inputs that shape model behaviour so the model follows an unintended instruction path. In practice, the target is often the model’s trust in what it reads, not the model’s base capability.
There is no single industry consensus label for every variant, so practitioners often use adjacent terms such as prompt injection, indirect prompt injection, or context poisoning depending on whether the abuse occurs through user text, retrieved content, or retained memory. The common failure pattern is the same: the system accepts context without enough source, intent, or scope validation.
For a control-oriented reference point, NIST SP 800-53 Rev 5 Security and Privacy Controls is useful for mapping the surrounding governance and access-control expectations, even though the term itself is more specific to AI context abuse.
Examples and Use Cases
- A retrieved support article contains hidden or conflicting instructions that cause the model to ignore the user’s request and follow the embedded text instead.
- A long-running chat session accumulates memory entries that later override the current user’s intent, especially when the system treats stored memory as always applicable.
- An agent reads tool output or web content that has been shaped to steer its next action, creating a mismatch between the visible task and the hidden instruction source.
- A multimodal workflow ingests documents with subtle instruction fragments, and the model follows them because the retrieval layer does not separate content from control signals.
- A product team allows broad context reuse for convenience, but the trade-off is that stale or irrelevant context can persist and steer later outputs in ways users did not expect.
The practical distinction is whether the system treats all context as equally trusted. Context-jacking becomes more likely when retrieval, memory, and conversation history are merged without clear provenance or scope rules.
Security Implications
When context-jacking succeeds, the model may produce unsafe, misleading, or policy-violating output while appearing to obey normal instructions. That can degrade answer integrity, break escalation logic, or trigger unauthorized tool use in agentic workflows.
The most serious failure mode is not only incorrect text generation. It is the redirection of downstream actions: the model may summarise the wrong source, expose data from an untrusted document, or follow an attacker-shaped instruction embedded inside retrieved context. In systems that chain model output into automation, that can become a business logic failure with real operational consequences.
Observable symptoms often include sudden instruction switching, unexplained refusals, unusual tool calls, or outputs that mirror source text too closely. Practitioners should treat those as signals that context is being trusted more than it is being validated. The blast radius grows when the same context is reused across sessions, users, or tools.
Domain and Governance Relevance
Context-jacking matters most in AI systems that combine retrieval, memory, and action. In that setting, the security problem is not just content pollution but control-plane pollution: the model’s decision path can be altered by inputs that were never meant to function as policy.
For NHI and agentic AI environments, the stakes rise because the model may act with delegated authority. If a tool-using agent consumes poisoned context, the resulting behaviour can affect service accounts, APIs, workflows, and approvals that sit outside the chat itself. That makes provenance, scope separation, and context lifecycle governance part of the trust model, not an optional hardening layer.
In practice, this term sits at the intersection of AI security, identity-bound action, and content integrity. The governing question is whether the system can distinguish user intent from surrounding data before that data influences a decision or executes an action.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Agentic AI Top 10 and MITRE ATLAS address the attack and risk surface, while NIST AI RMF, NIST AI 600-1 and CIS Controls v8 set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| NIST AI RMF | Map — Map | Context-jacking undermines AI system risk mapping and trust assumptions. |
| Recommendation — Map context sources and trust boundaries to identify where untrusted content can steer model behavior. | ||
| NIST AI 600-1 | GOV — Govern | AI governance must cover context provenance, memory, and instruction handling. |
| Recommendation — Govern context ingestion rules so hidden or stale instructions cannot override intended policy. | ||
| OWASP Agentic AI Top 10 | A1 — Input Validation and Sanitization | Agentic systems must validate external context before it can influence actions. |
| A4 — Tool and Action Authorization | Poisoned context can redirect tool use or delegated actions. | |
| Recommendation — Validate retrieved and conversational inputs before they reach agent decision logic. Authorize each tool action independently of model-generated context. | ||
| MITRE ATLAS | AML.TA0001 — Reconnaissance | Attackers probe context handling to discover how to steer model behavior. |
| Recommendation — Hunt for probing patterns that test how context alters model responses. | ||
| CIS Controls v8 | 6 — Access Control Management | Context-jacking often exploits overbroad access to retained or retrieved context. |
| Recommendation — Restrict who and what can modify reusable context, memory, and retrieval sources. | ||
Related resources from NHI Mgmt Group
- Why do traditional network firewalls and application controls fail against prompt injection and context-jacking?
- What is the Model Context Protocol (MCP) and why does it matter for security?
- What is MCP in the context of AI security?
- When is it appropriate to implement MCP in the context of AI systems?
Deepen Your Knowledge
Reviewed and updated by the NHIMG editorial team on September 7, 2026.
NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org