Join our Newsletter — 33% off our NHI Course

Notifications
Clear all

MITRE ATT&CK label prediction from Sigma rules: what are teams missing?


(@nhi-mgmt-group)
Member Moderator
Joined: 1 year ago
Posts: 17031
Topic starter  

TL;DR: Among 2,733 Sigma detection rules, Claude Opus 4.5 reached the highest hierarchical F1 score at 66%, while Gemini 3.0 Flash achieved the best recall at 71% and open-weight DeepSeek v3.2 led its class at 48% F1, according to Cotool. The result is a reminder that detection engineering workloads still depend on accurate adversary-technique mapping, not just model fluency.

NHIMG editorial — based on content published by Cotool: Sigma Detection Classification Jan 2026

By the numbers:

Questions worth separating out

Q: How should security teams use LLMs to map Sigma rules to MITRE ATT&CK?

A: Use them as enrichment assistants, not as final authorities.

Q: Why do ATT&CK labels from Sigma rules often need human review?

A: Because Sigma expresses detection logic, while ATT&CK expresses adversary behaviour.

Q: What breaks when ATT&CK technique mapping is inconsistent across detections?

A: Inconsistent mapping weakens hunting, distorts reporting, and makes it harder to compare detections across teams or environments.

Practitioner guidance

  • Use AI for first-pass ATT&CK enrichment only Route model output into analyst review queues rather than auto-publishing labels into SIEM content, because hierarchical matches can still hide technique-level mistakes.
  • Measure correction rates on your own Sigma corpus Track how often analysts change predicted technique IDs across your highest-value detections, then use that error pattern to decide whether the model is fit for enrichment.
  • Separate parent-technique and sub-technique workflows Apply different review thresholds when the model predicts a parent ATT&CK technique versus a sub-technique, since the operational risk of over- or under-specific mapping is not the same.

What's in the full report

Cotool's full analysis covers the benchmark methodology and model-by-model results this post intentionally leaves at a higher level:

  • Detailed F1, precision, recall, cost, and latency results for all 12 evaluated models
  • The benchmark scoring method, including hierarchical credit for parent and sub-technique predictions
  • Sample Sigma rule inputs showing how labels were inferred from detection logic
  • Model recommendation notes explaining when higher recall may be more useful than exact precision

👉 Read Cotool's benchmark analysis of LLM performance on Sigma-to-ATT&CK mapping →

MITRE ATT&CK label prediction from Sigma rules: what are teams missing?

Explore further

View Full Forum →  |  NHI Foundation Course →



   
Quote
(@mr-nhi)
Member Moderator
Joined: 3 months ago
Posts: 16618
 

Technique mapping is becoming a governance layer, not just a classification task. When teams use LLMs to infer ATT&CK techniques from Sigma rules, they are not merely automating annotation. They are influencing how detections are prioritised, how hunts are scoped, and how control gaps are reported to leadership. That makes technique mapping part of operational governance, especially when SIEM and SOAR workflows depend on consistent labels. The practitioner conclusion is simple: treat ATT&CK enrichment as controlled metadata, not free-form AI output.

A question worth separating out:

Q: How do teams know if ATT&CK enrichment is actually helping detection engineering?

A: Measure analyst correction rates, coverage of important technique families, and whether enriched rules improve triage speed or hunt precision. If the model adds noise without improving review quality, it is not helping. The useful signal is not perfect accuracy, but whether it makes detection content more consistent and more actionable.

👉 Read our full editorial: Sigma rules and MITRE ATT&CK labels still challenge LLMs



   
ReplyQuote
Share: