Join our Newsletter — 33% off our NHI Course

How should security teams decide where to red team an LLM application?

Start with the workflow boundaries that create real authority: retrieval, external tools, connectors, and any action the model can trigger. Then test how prompts can influence those paths. If the application includes agents or multi-step orchestration, add chained scenarios that exercise the full path from input to side effect.

Where to red team the LLM application first

Start with the places where the model can cross a trust boundary or trigger a side effect. That usually means retrieval, tool calls, connectors, workflow orchestration, and any path that can read, write, send, approve, or retrieve data outside the prompt. Those are the highest-value red team targets because they convert model influence into real action.

For practical test design, permission-aware retrieval and connector handling deserve early attention because they often decide whether the model can surface data it should not see. The same is true for enterprise AI copilots, where the meaningful risk is usually not the chat text itself but the attached search, mailbox, document, or ticketing actions.

If the application is mostly conversational with no external actions, the red team scope should stay narrower and focus on prompt influence, context leakage, and instruction hierarchy failures. If it is agentic, the scope expands to the full chain: input, memory, retrieval, tool selection, authorization, and the final effect in the downstream system.

What makes a red team target worth testing

The best targets are the components that change the model from a language interface into an operational actor. A harmless model answer is low value; a model answer that can retrieve restricted content, create tickets, send messages, call APIs, or alter records is where security teams learn something important about blast radius and control failure.

That is why agent behavior should be tested around authority boundaries, not only around model quality. Red teaming AI agents for identity abuse is especially useful when the application has approvals, delegation, shared credentials, or tool permissions that can be steered by malformed instructions. The red team objective is to see whether the system honors the intended principal, not just whether the model produces the right text.

Workflow mapping should also separate display risk from execution risk. Some paths only expose sensitive data to the user, while others actually create a side effect in another system. Those two cases require different tests, different evidence, and different severity thresholds.

How to cover chained and multi-step failures

Once the application has agents, planners, or multi-step orchestration, single-turn prompt attacks are no longer enough. Red teams should build chained scenarios that move from benign-looking input to retrieval, then tool use, then an action that changes state. The value is in testing whether the system can be led across each step without proper revalidation.

Agentic AI security guidance is relevant here because it frames the practical attack surface as inputs, memory, tools, and orchestration rather than only prompts. In a mature test plan, that means exercising cross-step abuse such as poisoned context, tool escalation, bad handoffs between sub-agents, and actions that proceed after the original instruction is no longer trustworthy.

Teams should also test the failure mode where the model appears to comply correctly at each individual step but still produces an unsafe end result. That is common in systems that split work across planners, retrievers, and executors, because local safety checks can miss global unsafe intent.

Risk and Threat Considerations

LLM applications are most exposed when prompt influence reaches a real entitlement, connector, or execution path. Attackers do not need perfect jailbreaks if they can steer the system into retrieving restricted data, misusing a tool, or taking an action on their behalf.

Failure mechanism: The model is induced to trust malicious instructions, over-broad context, or weakly governed connectors, then propagates that influence into retrieval or a side effect. In agentic systems, the failure often compounds across steps, because one unsafe decision can unlock the next.

Impact: The result can be data exposure, unauthorized actions, integrity loss, or abuse of connected systems. The more the application can do beyond answering questions, the more important it is to red team the boundaries where it can act.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

OWASP Agentic AI Top 10 addresses the attack and risk surface, while NIST AI RMF and OWASP ASVS set the governance and control requirements practitioners need to meet.

Framework Control / Reference Relevance
NIST AI RMF Govern LLM red teaming needs governance over model risk, testing scope, and deployment decisions.
Recommendation — Define red-team scope, thresholds, and sign-off criteria for LLM workflows before release.
OWASP Agentic AI Top 10 ASI02 — Tool Misuse Tool-calling paths are the highest-value red-team targets for agentic LLM applications.
ASI03 — Identity & Privilege Abuse Red teams should test whether the system can be steered into exceeding intended authority.
ASI08 — Cascading Failures Multi-step orchestration can turn one unsafe decision into a broader harmful chain.
Recommendation — Test whether prompts can drive unsafe tool calls or unauthorized downstream actions. Validate that agents cannot escalate privileges or act outside approved authority. Exercise chained scenarios to see whether failures compound across agent steps.
OWASP ASVS V8 — Authorization Any action-taking LLM workflow needs authorization boundaries around sensitive operations.
Recommendation — Verify that sensitive actions require enforced authorization outside model output.

Practitioner Guidance

What to prioritise: Put the first test budget into the paths that can cause irreversible or externally visible effects. Retrieval-only risks matter, but tool use, approval flows, and connectors usually justify earlier and deeper testing because they define real blast radius.

Decision rule: If a workflow can read, write, send, approve, or trigger anything outside the LLM boundary, test it as an operational control path, not just a prompt path. If it cannot, focus on instruction-following failures, leakage, and context handling.

What to verify: Each red team scenario should prove whether the system re-checks authority at the moment of action, not only at the moment of prompt ingestion. That is the difference between a language risk and a business risk.

Practitioner takeaway: Red team the LLM where it becomes capable of changing state, because that is where prompt weakness turns into concrete security exposure.