Join our Newsletter — 33% off our NHI Course

Notifications
Clear all

Blue team CTF log analysis benchmarks: what do the results mean?


(@nhi-mgmt-group)
Member Moderator
Joined: 1 year ago
Posts: 17031
Topic starter  

TL;DR: Frontier models still vary widely on blue team CTF work, with GPT-5.2 reaching about 69% accuracy, GPT-5.1 and Claude Opus 4.5 at 65%, and open-weight models trailing far behind, according to Cotool. The result is a reminder that SOC-grade investigation still depends on reliable tool use, context handling, and workflow discipline, not raw model capability alone.

NHIMG editorial — based on content published by Cotool: LLMjacking: How Attackers Hijack AI Using Compromised NHIs

By the numbers:

Questions worth separating out

Q: How should security teams govern AI SOC agents that use SIEM and EDR tools?

A: They should treat AI SOC agents as controlled investigative systems, not generic automation.

Q: Why do benchmark scores not fully capture SOC readiness for AI agents?

A: Because incident work depends on more than answer accuracy.

Q: What breaks when AI agents are given broad standing access?

A: Broad standing access breaks governance because the agent can move from one task to another without a fresh authorization check.

Practitioner guidance

  • Define agent-specific access boundaries Issue distinct identities for AI investigation agents, limit them to the log sources and actions they actually need, and revoke access automatically when the task ends.
  • Measure investigation quality beyond accuracy Track task completion, query efficiency, latency, and escalation frequency alongside benchmark accuracy so you can see whether the agent is dependable under SOC conditions.
  • Require audit-ready tool activity Log every search, dataset listing, and source-type description request from AI agents so analysts can reconstruct what the system saw and how it reached its answer.

What's in the full report

Cotool's full benchmark write-up covers the operational detail this post intentionally leaves for the source:

  • Per-model accuracy, cost, latency, and completion data across all 15 tested models for side-by-side evaluation
  • Methodology notes on the Splunk BOTSv3 environment, including tool access and task construction
  • Scenario-level examples that show how the benchmark questions map to real incident response workflows
  • Benchmark caveats that explain how human context was added or removed from the evaluation set

👉 Read Cotool's benchmark analysis of AI agent performance in blue team CTF scenarios →

Blue team CTF log analysis benchmarks: what do the results mean?

Explore further

View Full Forum →  |  NHI Foundation Course →



   
Quote
(@mr-nhi)
Member Moderator
Joined: 3 months ago
Posts: 16618
 

Agentic SOC benchmarking is now an identity governance problem, not just an AI accuracy problem. Once a model can query Splunk, it is operating as a non-human identity with tool access, decision latency, and audit requirements. That means SOC teams must govern permissions, logging, and revocation with the same seriousness they apply to service accounts. The practical conclusion is simple: if an AI can act inside your detection stack, it needs identity controls, not just model evaluation.

A question worth separating out:

Q: How do you know if AI-assisted investigations are actually working?

A: Look for defensible closure, not just shorter handling time. A working system should consistently correlate evidence from independent sources, reduce reopen rates, and produce conclusions that analysts trust enough to act on. If cases are closed quickly but frequently retriggered or manually corrected, the AI is speeding up uncertainty rather than resolving it.

👉 Read our full editorial: Blue team CTF benchmarking shows agentic SOC tasks remain hard



   
ReplyQuote
Share: