Join our Newsletter — 33% off our NHI Course

Notifications
Clear all

LLM triage accuracy varies by use case, so what should teams test first?


(@nhi-mgmt-group)
Member Moderator
Joined: 1 year ago
Posts: 20026
Topic starter  

TL;DR: A benchmark of 163 real-world security triage decisions found Gemini 3 Pro performed best overall at 74.8%, while the strongest model changed by use case, with different winners for phishing, account takeover, and network investigations, according to Legion AI. The result is a reminder that security teams need formal evaluation against their own workflows, not a one-model strategy.

NHIMG editorial — based on content published by Legion AI: Benchmarking Large Language Models for Automated Security Triage

By the numbers:

Questions worth separating out

Q: How should security teams evaluate LLMs for triage decisions?

A: Evaluate LLMs against the exact workflows, evidence types, and decision labels they will see in production.

Q: Why do different LLMs create different security risks for the same application?

A: Different models vary in training data, alignment, context capacity, and resistance to prompt manipulation, so the same workflow can behave differently across systems.

Q: What makes automated security triage fail in practice?

A: It fails when the workflow feeds the model incomplete or noisy evidence, when decision options are poorly defined, or when analysts assume one model fits every queue.

Practitioner guidance

  • Benchmark by queue, not by headline score Test phishing, account takeover, and network triage separately against your own investigation patterns, because the best model for one queue may underperform in another.
  • Audit workflow completeness before model tuning Check whether missing steps, partial evidence, or interrupted workflows are inflating apparent model error.
  • Route identity-heavy cases through stricter review gates Apply more conservative thresholds to phishing and account takeover cases where identity signals, access context, and user behaviour are tightly coupled.

What's in the full article

Legion AI's full benchmark covers the methodology and evaluation detail this post intentionally leaves for the source:

  • The dataset construction and cleaning rules used to reduce mock runs, missing data, and interrupted workflows.
  • The full confusion matrix and per-category performance results for each evaluated model.
  • The customer-environment tool stacks and workflow categories that shaped the triage decisions.
  • The benchmark's reasoning tags and annotation method for separating analyst disagreement from model error.

👉 Read Legion AI's benchmark of LLM performance across security triage use cases →

LLM triage accuracy varies by use case, so what should teams test first?

Explore further

View Full Forum →  |  NHI Foundation Course →



   
Quote
(@mr-nhi)
Member Moderator
Joined: 4 months ago
Posts: 19617
 

Model benchmarking in security operations is really workflow benchmarking. The headline score matters less than whether the model can make the right decision at the right point in a live investigation. In SOC and IAM-adjacent triage, evidence quality, decision options, and contextual enrichment drive outcomes as much as model capability. Practitioners should treat evaluation as a test of the whole operating model, not a contest between model names.

A question worth separating out:

Q: How do teams know if LLM triage is actually working?

A: Teams should look for stable decision quality across queues, low rates of avoidable false positives, and consistent agreement between analyst judgement and model recommendations. If performance changes sharply by use case, the model is not universally reliable and routing rules need to be tightened.

👉 Read our full editorial: Benchmarking LLMs for security triage shows model choice is use-case specific



   
ReplyQuote
Share: