TL;DR: AI red teaming shifts security testing from static infrastructure to the model, prompt, plugin, and agent layers, uncovering prompt injection, data leakage, and broken access control risks in real-world GenAI use, according to Lasso Security. The underlying issue is that traditional security assumes stable system behaviour, while AI behaviour changes with context, inputs, and chained tool calls.
Editorial analysis by NHI Mgmt Group, based on content published by Lasso Security: “What is Red Teaming in AI? Types, Components & Best Practices”.
Key questions
Q: How should security teams test GenAI systems for prompt injection?
A: Test the full path, not just the chat box.
Q: Why do GenAI guardrails fail when plugins and APIs are connected?
A: Because the model can turn a conversational request into a tool action, and tool actions often inherit broader permissions than the user intended.
Q: What are the signs that AI red teaming is failing to catch model vulnerabilities?
A: Red teaming is likely missing issues when reports show shallow coverage, repeated success against the same attack family, or no meaningful separation between vulnerability categories.
Practitioner guidance
- Define the exact GenAI test scope List the models, prompts, plugins, APIs, retrieval sources, and autonomous workflows that can influence output or action.
- Red team indirect prompt injection paths Test hidden instructions in PDFs, webpages, emails, and API responses to see whether the model treats untrusted content as instruction-bearing context.
- Validate tool-call authorization boundaries Check that each plugin or API call is constrained to the minimum scope needed for the current workflow, and that over-permissioned access is not inherited from the model wrapper.
Bottom line: AI red teaming is most useful when it tests how GenAI systems behave under hostile prompts, poisoned content, and chained integrations, not just how they answer clean questions.
Explore further
View Full Forum → | NHI Foundation Course → | Our Services → | Read the full analysis →
AI red teaming is really identity testing for delegated machine behaviour. The article describes a world where models, copilots, plugins, and agent workflows can trigger actions beyond the intent of the original prompt. That is not just application security, because the useful unit of analysis is the delegated authority chain behind the model. Practitioners should read this as a governance signal that AI systems are now exercising identity-like behaviour in production.
A few things that frame the scale:
- 85% of organisations lack full visibility into third-party vendors connected via OAuth apps, according to The State of Non-Human Identity Security.
- Lack of credential rotation is cited as the top cause of NHI-related attacks by 45% of organisations, followed by inadequate monitoring and logging at 37% and over-privileged accounts at 37%.
A question worth separating out:
Q: Should organisations treat AI plugins like privileged access?
A: Yes. AI plugins and connected APIs can act on behalf of the model, so they should be governed as privileged access paths with narrow scopes, explicit approval boundaries, and continuous review. If a plugin can reach customer records or production systems, it belongs in the same governance conversation as PAM and NHI controls.
👉 Read our full editorial: AI red teaming exposes where GenAI guardrails fail in practice
AI red teaming is now an identity and authorization test, not just a model-safety exercise. The article shows that GenAI failures emerge where prompts, plugins, and agent workflows intersect with delegated access. That means the real question is who or what the model is allowed to act for, and under what scope. Practitioners should treat model testing as a control over runtime authority.
A question worth separating out:
Q: Should security teams treat autonomous agents differently from chatbots?
A: Yes. Chatbots mainly answer questions, while autonomous agents can choose tools and continue execution across steps, which makes access scope, approval boundaries, and re-entrancy part of the threat model. A chatbot failure is often confined to the response. An agent failure can propagate into linked systems and compound before a human intervenes.
👉 Read our full editorial: AI red teaming exposes where GenAI guardrails fail in practice