Join our Newsletter — 33% off our NHI Course
Home› Glossary› AI Security› Context-Jacking
AI Security

Context-Jacking

← Back to Glossary
By NHI Mgmt Group Updated September 7, 2026 Domain: AI Security

Context-jacking is the abuse of an AI model’s conversation or retrieval context to redirect its behaviour away from intended policy. Attackers exploit hidden instructions, embedded data, or manipulated memory to influence outputs and actions. The risk rises when systems trust context without validating source, intent, and scope.

Expanded Definition

Context-jacking describes a class of prompt and retrieval abuse where an AI system is influenced by surrounding context rather than by the user’s intended instruction. That context may include hidden prompts, retrieved documents, long-conversation history, tool outputs, or memory objects that the system treats as authoritative.

The boundary matters: context-jacking is not simply “bad prompting” and not every model error is an attack. The term is used when an actor deliberately manipulates the contextual inputs that shape model behaviour so the model follows an unintended instruction path. In practice, the target is often the model’s trust in what it reads, not the model’s base capability.

There is no single industry consensus label for every variant, so practitioners often use adjacent terms such as prompt injection, indirect prompt injection, or context poisoning depending on whether the abuse occurs through user text, retrieved content, or retained memory. The common failure pattern is the same: the system accepts context without enough source, intent, or scope validation.

For a control-oriented reference point, NIST SP 800-53 Rev 5 Security and Privacy Controls is useful for mapping the surrounding governance and access-control expectations, even though the term itself is more specific to AI context abuse.

Examples and Use Cases

  • A retrieved support article contains hidden or conflicting instructions that cause the model to ignore the user’s request and follow the embedded text instead.
  • A long-running chat session accumulates memory entries that later override the current user’s intent, especially when the system treats stored memory as always applicable.
  • An agent reads tool output or web content that has been shaped to steer its next action, creating a mismatch between the visible task and the hidden instruction source.
  • A multimodal workflow ingests documents with subtle instruction fragments, and the model follows them because the retrieval layer does not separate content from control signals.
  • A product team allows broad context reuse for convenience, but the trade-off is that stale or irrelevant context can persist and steer later outputs in ways users did not expect.

The practical distinction is whether the system treats all context as equally trusted. Context-jacking becomes more likely when retrieval, memory, and conversation history are merged without clear provenance or scope rules.

Security Implications

When context-jacking succeeds, the model may produce unsafe, misleading, or policy-violating output while appearing to obey normal instructions. That can degrade answer integrity, break escalation logic, or trigger unauthorized tool use in agentic workflows.

The most serious failure mode is not only incorrect text generation. It is the redirection of downstream actions: the model may summarise the wrong source, expose data from an untrusted document, or follow an attacker-shaped instruction embedded inside retrieved context. In systems that chain model output into automation, that can become a business logic failure with real operational consequences.

Observable symptoms often include sudden instruction switching, unexplained refusals, unusual tool calls, or outputs that mirror source text too closely. Practitioners should treat those as signals that context is being trusted more than it is being validated. The blast radius grows when the same context is reused across sessions, users, or tools.

Domain and Governance Relevance

Context-jacking matters most in AI systems that combine retrieval, memory, and action. In that setting, the security problem is not just content pollution but control-plane pollution: the model’s decision path can be altered by inputs that were never meant to function as policy.

For NHI and agentic AI environments, the stakes rise because the model may act with delegated authority. If a tool-using agent consumes poisoned context, the resulting behaviour can affect service accounts, APIs, workflows, and approvals that sit outside the chat itself. That makes provenance, scope separation, and context lifecycle governance part of the trust model, not an optional hardening layer.

In practice, this term sits at the intersection of AI security, identity-bound action, and content integrity. The governing question is whether the system can distinguish user intent from surrounding data before that data influences a decision or executes an action.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

OWASP Agentic AI Top 10 and MITRE ATLAS address the attack and risk surface, while NIST AI RMF, NIST AI 600-1 and CIS Controls v8 set the governance and control requirements practitioners need to meet.

FrameworkControl / ReferenceRelevance
NIST AI RMFMap — MapContext-jacking undermines AI system risk mapping and trust assumptions.
Recommendation — Map context sources and trust boundaries to identify where untrusted content can steer model behavior.
NIST AI 600-1GOV — GovernAI governance must cover context provenance, memory, and instruction handling.
Recommendation — Govern context ingestion rules so hidden or stale instructions cannot override intended policy.
OWASP Agentic AI Top 10A1 — Input Validation and SanitizationAgentic systems must validate external context before it can influence actions.
A4 — Tool and Action AuthorizationPoisoned context can redirect tool use or delegated actions.
Recommendation — Validate retrieved and conversational inputs before they reach agent decision logic. Authorize each tool action independently of model-generated context.
MITRE ATLASAML.TA0001 — ReconnaissanceAttackers probe context handling to discover how to steer model behavior.
Recommendation — Hunt for probing patterns that test how context alters model responses.
CIS Controls v86 — Access Control ManagementContext-jacking often exploits overbroad access to retained or retrieved context.
Recommendation — Restrict who and what can modify reusable context, memory, and retrieval sources.

Deepen Your Knowledge

Sign up to our weekly newsletter — get 33% off our NHI Foundation Level Course

    NHIMG Editorial Note
    Reviewed and updated by the NHIMG editorial team on September 7, 2026.
    NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org