Join our Newsletter — 33% off our NHI Course

How should security teams implement threat modeling for GenAI applications that use RAG, agents, and external inputs?

Teams should start by mapping all untrusted data flows, including prompts, retrieved documents, tool outputs, and user-supplied content. Then identify the assets that matter most, such as system prompts, memory, secrets, and tenant data. Prioritise impact-driven scenarios first, because GenAI failures often appear as data leakage, prompt injection, or unsafe tool use rather than classic perimeter compromise.

Model the GenAI application as a set of trust boundaries, not a single prompt

The right threat model starts by splitting the application into distinct data flows and decision points: the user prompt, retrieved content, agent planning, tool invocation, memory, and any external API or database call. That separation matters because the attack surface changes at each handoff. A retrieval layer can leak data, an agent can overreach, and a tool can turn a harmless query into an unsafe action.

For RAG systems, the key question is not just whether the model is accurate, but whether untrusted content can influence what the system sees, remembers, or does. For agentic workflows, the key question is which actions are authorized, which are merely possible, and which should require confirmation or tighter policy.

That is why Agentic AI Security Guide is useful here: it gives a layered threat model for inputs, memory, tools, orchestration, and identity, which mirrors how these systems fail in practice.

Prioritise the assets that create the highest blast radius

Not every component deserves equal attention. The most important assets are the ones whose compromise changes what the system can reveal or execute: system prompts, secrets, retrieval indexes, memory stores, tenant data, and privileged tool credentials. If an attacker can influence or extract one of those, the impact is usually much larger than a normal model error.

In practice, this means the threat model should ask which assets are reachable from untrusted input and which ones are merely downstream. A retrieved document is not the same as a secret, but if the document can steer the agent into exposing a secret or calling a privileged tool, the risk becomes direct. The same logic applies to memory, because cross-session or cross-user contamination turns a quality issue into a confidentiality issue.

Permission-Aware RAG Guide is a strong companion for this part of the model because it focuses on retrieval-time permissions, over-sharing, and the control points that prevent RAG from becoming a data exfiltration path.

Use impact-driven scenarios to test prompt injection, tool abuse, and leakage paths

GenAI threat modelling works best when it is scenario-led. Instead of asking only “what can go wrong?”, ask “what is the worst plausible outcome if this untrusted input is malicious or simply wrong?” That usually surfaces failures faster than abstract taxonomies. Common impact paths include leaking tenant data, overwriting memory, triggering unsafe tool calls, or making the assistant act on behalf of a user without enough checks.

For RAG, the scenario may be indirect prompt injection through retrieved text. For agents, it may be tool misuse or a bad instruction chain that causes the system to take an action outside its intent. For external inputs, it may be a vendor feed, web page, or file upload that looks benign but becomes execution guidance once the model consumes it.

When you want a structured technique for those scenarios, Threat Modelling AI Agents helps translate trust boundaries into concrete attack trees and impact paths. For a broader external reference, CSA MAESTRO agentic AI threat modeling framework is useful for mapping multi-agent orchestration, autonomy, and tool-use risk.

Risk and Threat Considerations

GenAI applications are especially vulnerable to trust abuse because they routinely consume untrusted text, untrusted documents, and tool output that looks legitimate but is not trustworthy. The practical risk is not just model hallucination, it is data leakage, instruction hijacking, and unsafe delegation that can move a failure from “bad answer” to “bad action”.

Failure mechanism: Malicious or malformed input shapes retrieval, memory, planning, or tool selection so the system follows attacker-controlled instructions or exposes data it should not reveal.

Impact: Organisations can lose tenant isolation, leak secrets or sensitive documents, and trigger privileged actions that were never intended by the user or the business process.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

OWASP Agentic AI Top 10 addresses the attack and risk surface, while NIST AI RMF and NIST SP 800-53 Rev 5 set the governance and control requirements practitioners need to meet.

Framework Control / Reference Relevance
NIST AI RMF Govern GenAI threat modeling needs AI risk governance and documented risk treatment.
Recommendation — Define AI risk owners and require threat models for high-impact GenAI workflows.
NIST SP 800-53 Rev 5 RA-3 — Risk Assessment Threat modeling is a risk assessment activity for GenAI trust boundaries and scenarios.
AC-6 — Least Privilege Agents and tools need constrained authority to limit unsafe actions.
SI-10 — Information Input Validation Untrusted prompts, documents, and tool outputs require validation and filtering.
Recommendation — Assess untrusted inputs, retrieval paths, memory, and tools as part of the system risk review. Limit each agent and tool to the minimum permissions required for its task. Validate external inputs before they can influence retrieval, memory, or actions.
OWASP Agentic AI Top 10 ASI01 — Agent Goal Hijack Prompt injection and instruction steering can redirect agent objectives.
ASI02 — Tool Misuse Threat modeling must cover unsafe or unintended tool invocation.
ASI03 — Identity & Privilege Abuse Agent authority and access can be overused or abused in GenAI workflows.
Recommendation — Model goal hijack scenarios where untrusted content changes the agent's intended task. Map every tool to an authorization rule and abuse scenario before deployment. Constrain agent privileges and require checks for high-impact actions.

Practitioner Guidance

What to prioritise: Start with the flows that combine untrusted input and authority, especially retrieval pipelines, tool-calling paths, and any memory that persists across users or sessions. Those are the places where a small input problem can become a security incident.

What to verify: Confirm that each tool call, retrieval step, and external connector has an explicit trust decision attached to it, with clear ownership for who can approve, block, or log that action. If you cannot explain who is responsible for the decision, the control is not ready.

Practitioner takeaway: Treat GenAI threat modelling as an exercise in controlling influence and authority, not just filtering prompts; the highest-value controls are the ones that stop untrusted content from becoming privileged action.