Prompt injection is a runtime attack that manipulates an AI system through crafted input or retrieved content. Model poisoning targets the training or fine-tuning data so the model learns harmful behavior or backdoors before deployment. Both matter, but they require different controls. Prompt injection calls for workflow isolation and tool restrictions, while poisoning needs data governance, provenance, and dataset vetting.
How runtime attacks differ from training-time attacks
Prompt injection and model poisoning attack different parts of the generative AI lifecycle. Prompt injection is usually a live interaction problem, where an attacker smuggles malicious instructions into a prompt, chat thread, retrieved document, or tool output. Model poisoning is a supply-chain and data-integrity problem, where the attacker corrupts training or fine-tuning material so the model internalises the attacker’s intent before it is deployed.
The distinction matters because the security boundary moves. With prompt injection, the model may be behaving exactly as trained, but the runtime context has been manipulated. With poisoning, the model itself is altered or biased in advance, so the failure can appear normal until the model is used at scale.
These attack paths also create different operational symptoms. Prompt injection often shows up as unexpected tool calls, scope violations, data leakage, or instruction-following drift during a session. Model poisoning is more likely to show up as biased completions, hidden backdoors, degraded policy adherence, or behaviour that only appears when specific trigger phrases or patterns are present.
Why the controls are different
Prompt injection is best handled by reducing what the model can do with untrusted input, especially when the system can browse, retrieve, act, or call tools. That means isolating workflows, narrowing tool permissions, validating retrieved content, and limiting how much authority the model can exercise from a single prompt chain. NHIMG’s Agentic AI Security Guide is useful here because it treats inputs, memory, tools, and orchestration as separate control planes.
Model poisoning needs upstream controls instead. The central question is whether the data used to train or tune the model is trustworthy, traceable, and reviewable. That is why provenance, dataset vetting, filtering, signing, and change control matter more than runtime guardrails alone. A poisoned corpus can survive even very strong prompting policies if the model has already learned the attacker’s behaviour.
In practice, both controls are needed in mature systems. A model that is protected at runtime but fed untrusted training data still inherits hidden behaviour, while a clean model that is exposed to untrusted live instructions can still be steered into unsafe actions.
How to tell which problem you are actually seeing
Use the point of failure to separate the two. If the malicious influence arrives through a user prompt, a retrieved page, an email, a document, or a tool response during execution, you are dealing with prompt injection. If the malicious influence was introduced earlier, through the dataset, corpus, labels, embeddings, or fine-tuning pipeline, the more likely issue is poisoning.
That distinction is especially important for investigation and containment. Runtime attacks are usually bounded to the current session, affected connector, or affected workflow. Poisoning is harder to unwind because the bad influence may be embedded in a model version, a training snapshot, or a reused dataset that has already spread across environments.
NHIMG’s AI Supply Chain Security and AI-BOM Guide helps frame poisoning as a provenance problem, while the Enterprise AI Copilot Security Guide shows why runtime oversharing and connector control remain necessary even when the model itself is trusted.
Risk and Threat Considerations
Both attack types are dangerous because they target different trust assumptions. Prompt injection exploits the assumption that runtime content is benign, while model poisoning exploits the assumption that training data is clean. In agentic systems, either path can lead to data exfiltration, unauthorised actions, or malicious behaviour that looks like a legitimate model output.
Failure mechanism: A model follows attacker-supplied instructions at runtime, or it learns attacker-supplied behaviour during training, so the organisation loses control over what the system says, retrieves, or does.
Impact: The result can be confidentiality loss, unsafe tool use, backdoored responses, corrupted decision support, or broad downstream trust failure across applications that reuse the model or its outputs.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Agentic AI Top 10 addresses the attack and risk surface, while NIST AI RMF and CIS Controls v8 set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP Agentic AI Top 10 | ASI01 — Agent Goal Hijack | Prompt injection can steer an agent away from intended goals at runtime. |
| ASI03 — Identity & Privilege Abuse | Prompt injection and poisoning can drive unsafe actions through over-extended agent authority. | |
| ASI04 — Agentic Supply Chain Vulnerabilities | Model poisoning is a supply-chain integrity problem in AI training and tuning pipelines. | |
| Recommendation — Constrain agent objectives and reject untrusted instructions that alter task intent. Limit agent privileges and separate high-impact actions from routine model output. Vet data, models, and dependencies before they enter training or deployment pipelines. | ||
| NIST AI RMF | Generative AI risk management profile | The subject is a core generative AI security comparison covering provenance and runtime misuse. |
| Recommendation — Align controls to separate runtime prompt risks from training-data integrity risks. | ||
| CIS Controls v8 | CIS-3 — Data Protection | Poisoning and prompt-driven leakage both depend on protecting sensitive AI inputs and corpora. |
| Recommendation — Protect training data, prompts, and retrieved content with access and integrity controls. | ||
Practitioner Guidance
What to prioritise: Treat prompt injection as a runtime containment problem and poisoning as a data provenance problem. If the system can take actions, lock down tool scope first; if the model is being trained or tuned, verify dataset lineage before expanding deployment.
What to verify: Confirm which layer was touched, prompt, retrieval, tool output, fine-tune set, or pretraining corpus. If you cannot trace the source of influence, you cannot choose the right control set or know whether the issue is isolated to one session or baked into the model version.
Practitioner takeaway: The strongest teams do not treat these as variations of the same bug; they separate runtime trust from training-data trust, because the right containment action depends on which layer the attacker actually controlled.
Related resources from NHI Mgmt Group
- What is the difference between prompt injection and jailbreaking in AI security?
- What is the difference between prompt injection and data poisoning in LLM security?
- What is the difference between prompt injection and instruction override in AI security?
- What is the difference between prompt extraction and prompt injection in enterprise AI security?