TL;DR: GPT-5.1 and Claude Opus 4.5 tied at 65% accuracy on BOTSv3 SecOps tasks, according to Cotool’s benchmark update, while Opus 4.5 finished in 122 seconds on average and GPT-5.1 delivered top-tier accuracy at about $1.67 per task. The result is a clearer enterprise signal: model choice for security operations should be driven by task efficiency, not raw latency alone.
At a glance
What this is: Cotool’s benchmark update shows that GPT-5.1 and Claude Opus 4.5 now lead on SecOps task accuracy, with Opus 4.5 standing out for speed.
Why it matters: For IAM, SOC, and security automation teams, the finding matters because agent selection affects investigation quality, cost, and control over long-horizon workflows that increasingly intersect with identity and access data.
By the numbers:
- GPT-5.1 and Claude Opus 4.5 tied for the highest overall accuracy at 65% on the BOTSv3 benchmark.
- Claude Opus 4.5 completed benchmark tasks in 122 seconds on average, faster than the other newly tested models.
- GPT-5.1 achieved top-tier accuracy at roughly $1.67 per task in Cotool’s cost comparison.
👉 Read Cotool's benchmark update on GPT-5.1, Claude Opus 4.5, and Gemini 3 Pro in SecOps
Context
AI models used in security operations are no longer judged only by raw accuracy. In long-horizon SecOps workflows, the practical questions are how quickly a model converges, how many tool calls it burns, and whether it can sustain investigation quality across noisy log data and multi-step reasoning. That is especially relevant when the workflow touches identity-adjacent evidence such as cloud access, tokens, and account activity.
Cotool’s benchmark update uses the Splunk BOTSv3 dataset, which is built around realistic blue-team investigation tasks. The article’s core finding is that the frontier is narrowing for some models, but the strongest choices are not identical on speed, cost, and completion reliability. For practitioners, that means model evaluation has to move from abstract capability claims to workload-specific control over SecOps automation.
This is a typical pressure point for security teams experimenting with AI assistants in investigation pipelines: the model that looks best on paper may not be the one that behaves best under operational constraints.
Key questions
Q: How should security teams evaluate AI models for defensive cyber work?
A: Use representative tasks, not generic prompts. Score the model on accuracy, runtime cost, latency, and completion reliability against the actual workflows you expect it to support, then decide whether the result is good enough for triage, investigation, or analyst assistance. A benchmark only matters if it reflects your operational boundary.
Q: Why does reasoning efficiency matter in security operations?
A: Reasoning efficiency matters because security investigations often require multiple evidence steps, not a single answer. A model that reaches the right conclusion with fewer tool calls can reduce cost, analyst friction, and workflow drift. In practice, efficiency becomes a control signal for whether the model can sustain operational use.
Q: What breaks when AI models are chosen only on raw speed?
A: Teams can end up with fast models that produce incomplete investigations, weak evidence chaining, or poor convergence on complex log data. That can increase analyst workload and create false confidence in automation. Speed matters, but only when it does not reduce task completion or decision quality.
Q: How should organisations assign different AI models to different SecOps tasks?
A: Use lower-latency models for enrichment and routine triage, and reserve stronger reasoning models for forensic work or multi-step investigations. That segmentation matches model capability to task difficulty and keeps cost aligned with business need. A single-model strategy often wastes both budget and analyst time.
Technical breakdown
Why SecOps benchmarks depend on long-horizon reasoning
Security operations benchmarks are different from simple Q&A tests because the model must inspect logs, form hypotheses, call tools, and revise its path over several turns. That makes long-horizon reasoning central to performance. In practice, a model can score well on isolated tasks yet still struggle when the investigation requires persistence, context retention, and disciplined tool use across a noisy evidence set like BOTSv3.
Practical implication: benchmark models against the actual investigation patterns your team runs, not against generic chat performance.
Tool-call efficiency and task completion are different controls
A model that uses fewer tool calls is not automatically better, but tool efficiency can signal clearer reasoning and lower operational drag. Completion rate is a separate control signal because it shows whether the model can finish the task at all. In an operational SOC, both matter: a fast partial answer can be less useful than a slower but completed investigation, especially when the workflow depends on evidence chaining and analyst review.
Practical implication: measure completion rate, tool-call count, and time-to-answer together before allowing a model into production workflows.
Cost, latency, and quality form a SecOps Pareto frontier
The article shows a classic tradeoff pattern: some models deliver strong accuracy at lower per-task cost, while others match accuracy but spend more compute or wall-clock time. That is a Pareto frontier problem, not a single best-model problem. For security teams, the right decision depends on whether the use case is alert enrichment, incident triage, or deep forensic investigation. The wrong optimisation target can create either overspend or analyst bottlenecks.
Practical implication: assign different models to different SecOps tiers instead of trying to standardise everything on one metric.
NHI Mgmt Group analysis
Model benchmarking for security operations is becoming a governance problem, not just an engineering exercise. Once an AI system is trusted to inspect logs, correlate evidence, and recommend actions, the question shifts from model capability to operational accountability. Security teams need evidence that a model can sustain reasoning under real SecOps conditions, especially where identity signals, cloud telemetry, and access anomalies intersect. The benchmark format matters because it creates a repeatable way to judge whether an AI assistant is fit for controlled use.
Reasoning efficiency is now a first-class selection criterion for AI-enabled SecOps. Cotool’s results suggest that a model can be more useful because it reaches an answer in fewer steps, not because it merely has lower inference latency. That distinction matters for security work where each extra tool turn increases cost, drift, and analyst friction. Practitioner takeaway: evaluate the number of decision cycles a model needs to resolve an investigation, not just its raw speed.
Task accuracy alone can hide brittle operational behaviour. The gap between completion rate, tool efficiency, and final accuracy shows why SecOps benchmarking must look at the whole workflow. A model that is accurate but slow may still be unsuitable for triage, while a model that is fast but incomplete may create false confidence. For AI governance, the control objective is not just best score, but predictable performance under workload pressure.
SecOps teams should treat model choice as workload segmentation. The article reinforces a broader pattern: different models fit different tiers of work. High-confidence investigation tasks, interactive triage, and low-latency enrichment do not require the same capability mix. The practical conclusion is to map models to task classes, then validate them against actual security operating procedures before allowing automation to expand.
Identity and access evidence deserves explicit attention in AI-driven investigation pipelines. Many SecOps use cases increasingly depend on identity-adjacent signals such as cloud logins, token activity, and privileged account behaviour. That makes the model selection question relevant to IAM and PAM as well as SOC operations. Practitioner takeaway: if the workflow touches access evidence, benchmark the model against identity-heavy scenarios rather than generic log analysis.
What this signals
AI model evaluation is becoming part of security control design. Teams that plan to use AI in investigations need a benchmark discipline that looks at accuracy, completion, and cost together. That is especially true when the workflow includes identity evidence, because poor model behaviour can distort access-event analysis before a human sees the result.
Reasoning efficiency will increasingly shape how security teams operationalise AI assistants. The practical signal is that model choice will fragment by use case, with different systems serving triage, enrichment, and deep investigation. Security leaders should expect governance questions about where a model may act, how much evidence it may touch, and what escalation thresholds apply before automation is trusted.
Identity-heavy investigations need their own test harness. If the SecOps workflow depends on authentication events, privileged access, or cloud account behaviour, the benchmark must include those scenarios explicitly. For practitioners, the next step is to align AI evaluation with identity telemetry and the access controls that govern it.
For practitioners
- Benchmark models against your real SecOps task mix Replicate the kinds of investigations your analysts actually perform, including cloud incidents, account abuse, and long-context log review. Use the same toolchain and scoring rules so model comparisons reflect operational reality rather than abstract benchmark performance.
- Measure completion rate alongside accuracy Track whether a model finishes the task, how many tool calls it needs, and how often it converges without analyst rescue. Completion reliability is especially important in alert triage and investigative workflows where partial answers create extra toil.
- Segment models by investigation tier Use faster models for enrichment and first-pass triage, then reserve higher-reasoning models for deep investigations that need multi-step evidence chaining. This reduces cost while keeping the hardest cases on the most capable systems.
- Validate identity-heavy scenarios separately Test models on authentication anomalies, privileged access events, and access-log correlation before approving them for security automation. Identity-driven investigations often expose reasoning gaps that generic benchmark tasks do not reveal.
Key takeaways
- The benchmark shows that AI model choice for SecOps is now a workflow problem, not a simple accuracy contest.
- GPT-5.1 and Claude Opus 4.5 lead on accuracy, but Opus 4.5 stands out for speed while GPT-5.1 offers stronger cost efficiency.
- Security teams should benchmark AI assistants against real investigation tasks, then segment models by the operational tier they can safely handle.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
MITRE ATT&CK address the attack and risk surface, while NIST AI RMF, NIST CSF 2.0 and NIST SP 800-53 Rev 5 set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| NIST AI RMF | MEASURE | The article is about evaluating AI model performance and operational reliability. |
| NIST CSF 2.0 | GV.RM-01 | Model choice affects risk management for security operations workflows. |
| NIST SP 800-53 Rev 5 | SI-4 | Security monitoring workflows depend on reliable analysis of security telemetry. |
| MITRE ATT&CK | TA0007 , Discovery; TA0009 , Collection | The benchmark centers on investigation of log evidence and attack traces. |
Use ATT&CK to structure investigation tasks and validate model coverage of discovery and collection.
Key terms
- Long-horizon cyber reasoning: The ability to investigate, remember, and re-plan across many steps while pursuing a security objective. In practice, this means the system can carry state across domains such as identity, cloud, and runtime, then use that state to continue toward the target after failures or dead ends.
- Tool Efficiency: A measure of how effectively a model uses external tools, such as search, queries, or analyzers, to reach an outcome. In security operations, lower tool usage can mean clearer reasoning, but only if the task is still completed accurately and with enough evidence to support analyst review.
- Task Completion Rate: The percentage of benchmark tasks a model finishes successfully within the evaluation constraints. For security automation, completion rate is a critical indicator because partial investigations can create more manual work, delay incident handling, and produce misleading confidence in the model’s output.
- Pareto Frontier: The set of options where improving one outcome, such as speed or cost, would worsen another, such as accuracy. In AI operations, this concept helps teams choose a model that fits the use case rather than chasing a single best score across all dimensions.
What's in the full report
Cotool's full research covers the operational detail this post intentionally leaves for the source:
- Per-model benchmark tables for GPT-5.1, Claude Opus 4.5, Gemini 3 Pro, Sonnet 4.5, and Haiku 4.5 across accuracy, cost, completion, and time.
- Task-level methodology for the BOTSv3 blue-team CTF environment, including how the agent harness was configured for investigation workflows.
- Extended notes on tool-call efficiency and token consumption that help teams decide where a model fits in the SecOps pipeline.
- The authors' planned failure analysis for the longer-context cases where some models needed more convergence than expected.
Deepen your knowledge
The NHI Foundation Level course, the industry's only accredited NHI security programme, covers NHI governance, machine identity security, secrets management, and identity lifecycle control. It gives security practitioners a common foundation for connecting identity governance to broader AI and automation risk.
Published by the NHIMG editorial team on August 18, 2026.
NHI Mgmt Group — the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org