Join our Newsletter — 33% off our NHI Course

Autonomous red teaming for LLMs: what security teams need now

 

(@nhi-mgmt-group)
Member Moderator
Joined: 1 year ago
Posts: 20739
Topic starter  

TL;DR: Autonomous red teaming is being positioned as a way to stress-test LLMs and agentic systems before attackers do, while OWASP updates on excessive agency, system prompt leakage, and RAG weaknesses show why traditional red teaming leaves blind spots, according to Lasso Security and OWASP. The real issue is not just model testing, but the lack of clear ownership and benchmarks for systems that now change behaviour, expose prompts, and pull external data at runtime.

Editorial analysis by NHI Mgmt Group, based on content published by Lasso Security: “Strengthening LLM Security from the Get-Go: Autonomous Red Teaming in Action”.

Key questions

Q: What breaks when LLMs are red-teamed only with traditional manual methods?

A: Manual red teaming misses the combinations of prompts, retrieval paths, and tool interactions that emerge only at runtime.

Q: Why do excessive agency and prompt leakage matter for AI governance?

A: Excessive agency turns an LLM system into a governed actor that can affect downstream workflows, while prompt leakage exposes the instructions that shape those actions.

Q: What are the signs that RAG is becoming a security problem?

A: Warning signs include weak control over retrieval sources, unclear provenance for indexed content, and responses that reflect data the system should not have reached.

Practitioner guidance

  • Define an owner for AI security benchmarking Assign a named security owner for each LLM or GenAI deployment so red-team findings do not stall between model, platform, and application teams.
  • Test excessive agency as a permissions problem Map which AI actions are truly allowed at runtime, then verify those permissions against the workflows the system can initiate without human approval.
  • Inventory system prompts and guardrail logic Treat prompts, policy instructions, and hidden control logic as governed assets that must be reviewed, protected, and version-controlled.

Bottom line: Traditional red teaming leaves gaps when LLM behaviour shifts through prompts, retrieval, and tool use at runtime.

Explore further

View Full Forum →  |  NHI Foundation Course →  |  Our Services →  |  Read the full analysis →


This topic was modified 2 hours ago by NHI Mgmt Group

   
Quote
(@mr-nhi)
Member Moderator
Joined: 5 months ago
Posts: 21364
 

Autonomous red teaming is becoming the only credible way to test LLMs that change behavior at runtime. Manual red teaming still matters, but it was built for systems whose boundaries could be enumerated and retested on a schedule. LLMs with dynamic retrieval, prompt-driven behaviour, and tool access alter those boundaries continuously. The implication is that security assurance for AI systems now has to be continuous, not episodic.

A few things that frame the scale:

  • 98% of companies plan to deploy even more AI agents within the next 12 months, despite documented rogue behaviour in 80% of current deployments, according to AI Agents: The New Attack Surface report.
  • Only 52% of companies can track and audit the data their AI agents access, leaving 48% with a complete blind spot for compliance and breach investigation.

A question worth separating out:

Q: What should organisations do first when securing LLMs and AI agents?

A: Organisations should start by defining ownership, permitted actions, and the boundaries of model access. Before advanced controls, they need to know who is accountable for prompts, retrieval sources, and tool permissions. Without that baseline, testing and monitoring cannot be tied to a clear security decision.

👉 Read our full editorial: Autonomous red teaming exposes the LLM security benchmark gap



   
ReplyQuote
(@mr-nhi)
Member Moderator
Joined: 5 months ago
Posts: 21364
 

Autonomous red teaming is becoming the only credible way to test LLMs that change behavior at runtime. Manual red teaming still matters, but it was built for systems whose boundaries could be enumerated and retested on a schedule. LLMs with dynamic retrieval, prompt-driven behaviour, and tool access alter those boundaries continuously. The implication is that security assurance for AI systems now has to be continuous, not episodic.

A few things that frame the scale:

  • 98% of companies plan to deploy even more AI agents within the next 12 months, despite documented rogue behaviour in 80% of current deployments, according to AI Agents: The New Attack Surface report.
  • Only 52% of companies can track and audit the data their AI agents access, leaving 48% with a complete blind spot for compliance and breach investigation.

A question worth separating out:

Q: What should organisations do first when securing LLMs and AI agents?

A: Organisations should start by defining ownership, permitted actions, and the boundaries of model access. Before advanced controls, they need to know who is accountable for prompts, retrieval sources, and tool permissions. Without that baseline, testing and monitoring cannot be tied to a clear security decision.

👉 Read our full editorial: Autonomous red teaming exposes the LLM security benchmark gap



   
ReplyQuote
(@mr-nhi)
Member Moderator
Joined: 5 months ago
Posts: 21364
 

Autonomous red teaming fills a coverage gap, not a governance gap: Traditional red teaming was designed for bounded systems with stable interfaces and slower change cycles. LLM deployments evolve through prompts, retrieval sources, and application logic that shift faster than manual test programmes can follow. The implication is that benchmarking alone is insufficient unless ownership and remediation are also defined.

A few things that frame the scale:

  • AI-related credential leaks surged 81.5% year-over-year in 2025, with the surrounding AI infrastructure leaking 5x faster than core LLM providers, according to the State of Secrets Sprawl 2026.
  • Only 23% of IT leaders were very confident in their organisation's ability to manage security and governance for GenAI deployments, according to a 2025 Gartner survey of 360 IT leaders.

A question worth separating out:

Q: How should teams govern AI systems that can take actions as well as generate outputs?

A: Treat the agent as a governed actor, not just a model output stream. Require action-level logging, tool-call traceability, authorization boundaries, and approval gates before the system can write to records or invoke downstream tools. If an AI system can change state, its authority must be scoped, monitored, and revocable like any other privileged non-human identity.

👉 Read our full editorial: Autonomous red teaming exposes the LLM security benchmark gap


This post was modified 2 hours ago by NHI Mgmt Group

   
ReplyQuote
Share:

Free weekly newsletter

Subscribe to the NHI & AI Identity Journal

The latest on NHI and Agentic AI security – articles, research, breaches, news and events every week.

Bonus 33% off our NHI Course when you subscribe.