TL;DR: Agentic AI red teaming that stays at the API endpoint misses the context pipeline, tool calls, and output sinks that shape real attack paths, according to Pillar Security. Coverage has to shift from prompt testing to runtime discovery, because the security question is no longer whether a model answers badly, but whether the agent can be driven to misuse connected systems.
Editorial analysis by NHI Mgmt Group, based on content published by Pillar Security: “Agentic Red Teaming: Five Dimensions Your Testing Should Cover”.
Key questions
Q: How should security teams red team AI agents that use tools and memory?
A: Security teams should test AI agents through the same interface and runtime path production uses, then validate the tools, memory stores, and downstream sinks those agents can reach.
Q: Why do API-level red team tests miss agentic AI risks?
A: Because the API endpoint is not the whole system.
Q: What are the signs that an agentic AI red team is too narrow?
A: A red team is too narrow when it focuses mainly on prompt injection scores, jailbreaks, or single-turn harmful output tests.
Practitioner guidance
- Test through the production interface Run red team traffic through the same UI or workflow that users and attackers actually use, so context rewriting and sink behaviour are part of the test.
- Map tool reachability before adversarial input Inventory which tools, MCP servers, data stores, and downstream systems the agent can actually reach before you start prompt attacks.
- Include output sinks in scope Validate rendered markdown, links, embedded content, database writes, message routing, and any other downstream execution path the agent can trigger.
Bottom line: Agentic red teaming fails when it tests the model response but ignores the surrounding runtime, because the real attack surface includes context handling, tool access, and output sinks.
Explore further
View Full Forum → | NHI Foundation Course → | Our Services → | Read the full analysis →
The API is not the attack surface for agentic systems. Agentic red teaming that stops at the model endpoint is testing a different system from the one in production. The context pipeline, tool invocation layer, and output sinks are where the real attack path lives, and those layers can rewrite, redirect, or amplify the model's behaviour. Practitioners should treat API-only results as partial evidence, not program coverage.
A few things that frame the scale:
- 98% of companies plan to deploy even more AI agents within the next 12 months, despite documented rogue behaviour in 80% of current deployments, according to AI Agents: The New Attack Surface report.
- Only 52% of companies can track and audit the data their AI agents access, leaving 48% with a complete blind spot for compliance and breach investigation.
A question worth separating out:
Q: How do organisations decide whether agentic red teaming is actually working?
A: Organisations should judge agentic red teaming by coverage of runtime paths, not by the number of prompts tested. If the programme can map verified tools, permissions, data flows, and downstream actions into reproducible exploit chains, it is working. If it only produces response-level failures, it is still testing the wrong surface.
👉 Read our full editorial: Agentic AI red teaming fails when tests stop at the API
The API is not the attack surface for agentic systems. Agentic red teaming that stops at the model endpoint is testing a different system from the one in production. The context pipeline, tool invocation layer, and output sinks are where the real attack path lives, and those layers can rewrite, redirect, or amplify the model's behaviour. Practitioners should treat API-only results as partial evidence, not program coverage.
A few things that frame the scale:
- 98% of companies plan to deploy even more AI agents within the next 12 months, despite documented rogue behaviour in 80% of current deployments, according to AI Agents: The New Attack Surface report.
- Only 52% of companies can track and audit the data their AI agents access, leaving 48% with a complete blind spot for compliance and breach investigation.
A question worth separating out:
Q: How do organisations decide whether agentic red teaming is actually working?
A: Organisations should judge agentic red teaming by coverage of runtime paths, not by the number of prompts tested. If the programme can map verified tools, permissions, data flows, and downstream actions into reproducible exploit chains, it is working. If it only produces response-level failures, it is still testing the wrong surface.
👉 Read our full editorial: Agentic AI red teaming fails when tests stop at the API
API-only red teaming is a fidelity problem, not a coverage problem. The model endpoint is only one component in the security boundary, while agentic systems rewrite context before inference and process outputs after inference. A test that ignores those layers can still be technically sophisticated and still validate the wrong system. The practitioner conclusion is that runtime fidelity must outrank prompt volume in every agent red team programme.
A few things that frame the scale:
- 67% of organisations still rely heavily on static credentials despite the risks they pose to agentic AI deployments, according to the 2026 Infrastructure Identity Survey.
A question worth separating out:
Q: How do teams judge whether an agentic red team finding is actually serious?
A: A serious finding identifies the entry point, the tool pivot, the data or workflow reached, and the business impact. If the test cannot reproduce a path across the runtime stack, the result is usually a prompt artefact, not evidence of exploitable agent behaviour.
👉 Read our full editorial: Agentic AI red teaming fails when tests stop at the API