TL;DR: AI triage often repeats the same mistakes because it lacks memory, according to torq research based on a review of more than 1,000 analyst corrections across four customer environments, with confidence-based learning improving verdict alignment to 92% versus 78% for similarity-based approaches. The governance problem is not just model accuracy, but whether AI SOC decisions become consistent enough to trust and operationalise.
NHIMG editorial — based on content published by torq: AI SOC learning gaps show why alert triage still needs memory
By the numbers:
- Torq reviewed more than 1,000 real analyst corrections across four customer environments to understand why the AI and the analyst kept disagreeing.
- Reflex matched the analyst’s corrected verdict 92% of the time and held within about three points of that across folds.
Questions worth separating out
Q: How can teams tell whether AI triage is actually improving SOC operations?
A: Look for lower manual processing time, fewer duplicate reviews, shorter disposition cycles, and faster removal of related malicious messages.
Q: Why do AI triage tools become unreliable after feedback loops?
A: They often store feedback as prompts or lookups rather than as durable decision memory.
Q: What breaks when an AI SOC cannot remember prior decisions?
A: The same alert can be judged differently on different days, which creates inconsistency, repeated analyst corrections, and lower trust in automation.
Practitioner guidance
- Define confidence-based routing rules Set explicit thresholds for when the AI may auto-triage an alert, when it should escalate to an analyst, and when it should stay in review because the confidence score is too low to justify automation.
- Track verdict consistency across shifts Measure whether the same alert pattern gets the same outcome over time, including after analyst feedback and across different shifts, because drift in verdicts is an early warning that the learning layer is not stable.
- Re-test after every model update Require regression testing against a fixed set of prior analyst decisions whenever the underlying model changes, so prompt drift and model drift are both visible before production behaviour changes.
What's in the full article
Torq's full analysis covers the operational detail this post intentionally leaves for the source:
- The customer-by-customer comparison between prompt-based learning and stateful model training in analyst triage.
- The Reflex confidence scoring approach and how it changes routing between automation and human review.
- The measured accuracy differences across corrected verdicts, overall outcomes, and threat-confirmation calls.
- The operational explanation of how model updates affect calibration and consistency over time.
👉 Read Torq's analysis of AI SOC memory, learning, and analyst trust →
AI SOC triage without memory: what changes for SOC teams?
Explore further
Decision memory is becoming a governance requirement for AI SOCs. A system that can explain an alert but not remember how the organisation decided on similar cases is not yet ready for operational trust. The article shows that prompt updates and lookup methods can improve surface behaviour, but they do not create durable judgement. For SOC programmes, this shifts AI from a detection aid to a governed decision layer, which should be evaluated like any other operational control.
A question worth separating out:
Q: How do teams know if AI SOC learning is actually working?
A: Look for stable verdicts on repeated alert patterns, higher agreement with analyst corrections, and fewer unnecessary re-reviews after retraining. The key signal is not whether the model is active, but whether it produces consistent outcomes that match local policy across shifts and model updates.
👉 Read our full editorial: AI SOC learning gaps show why alert triage still needs memory