Security teams should scope AI penetration tests by identifying which AI pattern is in use, then testing the access paths it creates. Direct LLM calls, MCP servers, RAG pipelines, and agents each expose different trust boundaries, so the assessment should cover prompt injection, tool access, permission boundaries, input validation, and output handling across every layer the AI touches.
Scoping the Test by AI Pattern, Not by the Word “AI”
Penetration testing scope should start with the architecture pattern, because each pattern creates a different attack surface and trust boundary. A direct LLM integration is mainly about prompt handling and output safety; MCP adds protocol and tool authorization; RAG adds retrieval, indexing, and permission enforcement; agents add delegated action, chaining, and escalation risk. The test plan should reflect the component that can actually change security outcomes.
That means the scope document should name the exact flows under test: user input to model, model to retrieval layer, model to tool call, tool call to external system, and any human approval or policy gate in between. It should also state what is out of scope, such as model training or cloud infrastructure, unless those layers are part of the deployed application path.
For RAG-heavy systems, Permission-Aware RAG Guide is the right starting point when retrieval can surface data the caller should not see. For agentic systems, Agentic AI Security Guide helps frame the test around inputs, tools, memory, and orchestration rather than only the chat surface.
What Each Layer Should Be Tested For
Direct LLM calls should be tested for prompt injection, system prompt leakage, unsafe output handling, and reliance on user-supplied content that can change model behaviour. The important question is not whether the model can be tricked in the abstract, but whether a malicious prompt can change downstream decisions, expose secrets, or produce unsafe content that another system trusts.
MCP requires a different lens because it introduces a protocol layer and tool boundary. Test whether the server enforces caller identity, audience-bound tokens, least privilege, and correct separation between the model, the MCP client, and the underlying tool. In practice, a weak MCP scope often fails at authorization, not at generation quality.
RAG should be tested for retrieval poisoning, oversharing, document-level access control, vector-store exposure, and data contamination from untrusted sources. If the retrieval layer ignores user entitlements, the model may faithfully answer with data the caller was never allowed to access. Permission-Aware RAG Guide is especially useful for this layer because it treats permissions as part of retrieval, not an afterthought.
Agents should be tested for delegated authority, tool misuse, cross-step escalation, and the ability to combine individually safe actions into a harmful sequence. Where an agent can invoke tools, write files, send messages, or trigger workflows, the test must verify what happens when instructions conflict, when the context is poisoned, or when a tool response contains attacker-controlled data.
How to Design a Useful Test Plan for Mixed AI Apps
The best scope is usually a matrix: rows for data flows and components, columns for trust assumptions, permissions, and abuse cases. That makes it easier to see where the application crosses from user content into privileged action. It also prevents a common mistake, which is testing only the chat prompt while ignoring retrieval, middleware, and tool execution.
- Map each AI pattern to its trust boundary.
- List the privileged actions the pattern can trigger.
- Identify which secrets, tokens, or service identities it can reach.
- Define abuse cases for prompt injection, data leakage, and permission bypass.
- Test both direct abuse and chained abuse across layers.
If the system uses an agent, include tool authorization and human approval paths in scope, even when the underlying model is not directly exposed. If the system uses RAG, include retrieval permissions, index hygiene, and source trust. If the system uses MCP, include protocol-level access control and token handling before you spend time on model behaviour alone.
Risk and Threat Considerations
These systems fail when teams test the model as if it were the whole product. The real exposure is usually at the boundary between the model and the resources it can read or control. Prompt injection, over-broad retrieval, weak tool authorization, and secret exposure can each turn a seemingly benign assistant into a path for data theft or unauthorized action.
Failure mechanism: Attackers exploit whichever layer has the weakest boundary, then use the model’s trusted position to reach retrieval, tools, or privileged APIs. In mixed architectures, a flaw in one component can become an end-to-end compromise if the test does not follow the full execution path.
Impact: The result can be confidential data exposure, unauthorized business actions, credential theft, or agent-driven lateral movement into connected systems. In higher-trust environments, the blast radius can expand quickly because the AI component often sits close to search, collaboration, or automation workflows.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Non-Human Identity Top 10 and OWASP Agentic AI Top 10 address the attack and risk surface, while NIST SP 800-53 Rev 5 and OWASP ASVS set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP Non-Human Identity Top 10 | NHI-05 — Overprivileged NHI | Agents, tools and retrieval paths fail when access is broader than needed. |
| NHI-04 — Insecure Authentication | MCP and tool paths depend on caller authentication and token handling. | |
| NHI-02 — Secret Leakage | LLM, RAG and agent flows can expose API keys, prompts and retrieved data. | |
| Recommendation — Constrain AI-connected identities to least privilege and test for excess access. Verify authentication and token audience controls on every AI access path. Test for secret exposure in prompts, logs, retrieval stores and tool outputs. | ||
| OWASP Agentic AI Top 10 | ASI03 — Identity & Privilege Abuse | Agent tests must cover delegated authority and privilege escalation through tools. |
| ASI02 — Tool Misuse | MCP and agentic apps are defined by how tools are selected and invoked. | |
| ASI06 — Memory & Context Poisoning | RAG and agent memory can be manipulated to alter decisions or leak data. | |
| Recommendation — Red-team agent actions that can exceed intended identity or privilege boundaries. Validate that tools cannot be misused through prompt or context manipulation. Test memory and context for poisoning, cross-session leakage and unsafe reuse. | ||
| NIST SP 800-53 Rev 5 | AC-6 — Least Privilege | AI components should only access the data and tools required for the task. |
| IA-5 — Authenticator Management | MCP and tool integrations rely on secure credential lifecycle and handling. | |
| Recommendation — Apply least privilege to RAG stores, tools, and agent credentials. Rotate, protect and scope the credentials used by AI-connected services. | ||
| OWASP ASVS | V8 — Authorization | RAG and tool access need authorization checks beyond the model prompt. |
| V16 — Security Logging and Error Handling | AI abuse is hard to detect without logging inputs, outputs and tool use. | |
| Recommendation — Verify every AI-triggered action is authorized at the application boundary. Capture AI inputs, outputs and tool events needed for incident review. | ||
Practitioner Guidance
What to prioritise: Start with the highest-privilege path the application can take, not the most visible prompt. If the AI can read restricted data or invoke real actions, that path deserves first-pass testing before lower-value prompt-quality checks.
What to verify: Confirm that every retrieval source, tool, and external API is enforced by a real access decision, not by prompt instructions alone. Verify that the test can prove whether the application leaks data, changes state, or escalates privilege when given malicious input.
Common mistake: Teams often write one generic “LLM red team” scope and assume it covers MCP, RAG, and agents. It usually does not, because the failure mode shifts from text manipulation to authorization, delegation, and data-flow abuse as the architecture becomes more connected.
Practitioner takeaway: Scope the test around the action the AI can cause, because that is where the security boundary actually lives.
Related resources from NHI Mgmt Group
- How should security teams govern AI agents that use service accounts and MCP tools?
- How should security teams govern access when LLMs use MCP servers?
- How should security teams use audits and penetration tests together?
- How should security teams scope application penetration tests for modern cloud and AI-enabled systems?