TL;DR: Security teams are discovering that SOC staffing, tooling, and alert volumes no longer fit sustainable in-house operating models, according to Expel, with turnover, false positives, and 24x7 coverage costs compounding the problem. Data-driven sizing and hybrid coverage now matter more than aspirational staffing plans.
At a glance
What this is: This analysis argues that SOC economics, talent scarcity, and alert fatigue are pushing many organisations away from fully in-house security operations.
Why it matters: It matters to IAM and security practitioners because operational strain changes how teams detect identity abuse, triage privilege misuse, and sustain coverage across human and machine identities.
By the numbers:
- Women represent only 22% of cybersecurity professionals globally, constraining the talent pool before hiring even begins.
- A single analyst can handle 20 to 30 quality alerts per shift when investigations are thorough, but far more than that forces shallow triage.
- If analyst utilisation consistently exceeds 70%, the team is already operating in a burnout-prone zone rather than an efficient one.
👉 Read Expel's analysis of SOC staffing, alert fatigue, and operational cost
Context
Security operations depend on a balance between alert volume, analyst capacity, and tooling quality. When that balance breaks, detection becomes inconsistent, response slows, and organisations lose the ability to investigate the identity misuse and credential abuse that often sit inside noisy alert streams.
This article is fundamentally about SOC governance rather than a single tool choice. The identity connection is real because overworked teams are less able to spot suspicious privileged access, service account misuse, and other non-human identity behaviours that require careful review.
The starting position described here is increasingly common for mid-sized and larger organisations. The challenge is not whether to do security operations, but whether the traditional in-house model can still be sustained at the required level of coverage.
Key questions
Q: How should security teams size a SOC for sustainable coverage?
A: Start with real alert volume, average handling time, and the amount of time analysts can spend at full effectiveness before quality drops. Then model shift coverage, escalation, vacations, and turnover. If the math only works by assuming constant overtime, the SOC is under-sized and the control model is already fragile.
Q: Why do overloaded SOCs miss identity abuse and privileged access anomalies?
A: Because identity events often look normal until they are correlated with other signals. When the queue is full, analysts clear obvious alerts and lose time for deeper investigation, which means suspicious logins, service account misuse, and privilege changes can sit unreviewed long enough for attackers to progress.
Q: What do security teams get wrong about automated SOC reporting?
A: They often treat report generation as a formatting task instead of a control point. A useful generated report must reflect the actual timeline, evidence sources, and actions taken, or it becomes a polished summary with weak investigative value. The report should support handoffs, review, and auditability.
Q: Should organisations build, buy, or hybridise SOC operations?
A: The right answer depends on volume, talent access, and the need for continuous coverage. Build works when you can fund staffing, tuning, and retention over time. Buy or hybridise when the organisation needs immediate coverage, lower operational strain, or a better balance between internal expertise and managed triage.
Technical breakdown
SOC alert fatigue and why default detections fail
SOC alert fatigue happens when security tools generate more signals than analysts can meaningfully investigate. Default detection rules are usually tuned to avoid false negatives, which means they over-alert in real environments and swamp teams with noise. Without environment-specific tuning, the SOC spends time clearing queues instead of confirming incidents. The result is not just inefficiency. It creates blind spots where real attacks hide inside routine activity.
Practical implication: tune detections against your own environment and measure the false-positive rate before adding more tools.
Why 24x7 coverage has become economically fragile
Round-the-clock security operations require enough people to cover shifts, vacations, escalation, and specialist investigation, not just a minimum headcount on paper. Once turnover rises, each new hire also needs time to absorb internal context, which extends the period before the team becomes effective. The economics break when organisations assume that a small team can absorb growing volume without losing response quality. Sustainable coverage depends on workload, not intent.
Practical implication: size the SOC from actual handling time and escalation demand, not from a fixed staffing target.
How SOC operations intersect with identity governance
Identity activity is one of the most important signals in a SOC because attacker movement often shows up as abnormal login patterns, privilege changes, or abuse of service accounts. When analysts are overloaded, those identity cues are easier to miss, especially in environments with privileged access, automation tokens, and mixed human-machine workflows. That makes SOC capacity a governance issue for IAM and PAM teams as well as a monitoring problem.
Practical implication: align SOC triage workflows with IAM and PAM telemetry so identity abuse is investigated as a priority signal.
Threat narrative
Attacker objective: The attacker objective is to stay hidden long enough for detection to fail and response to arrive after meaningful damage has already occurred.
- Entry begins in the SOC workflow itself, where overwhelming alert volume and noisy defaults reduce the chance that malicious activity is noticed early.
- Escalation follows when analysts become unable to distinguish routine activity from real compromise, allowing identity abuse, persistence, or lateral movement signals to sit unreviewed in the queue.
- Impact occurs when critical incidents are discovered late, after attackers have already used the delay to extend access, exfiltrate data, or disrupt operations.
NHI Mgmt Group analysis
Alert volume is now a governance problem, not just an operations problem. When analysts cannot keep pace, the organisation has effectively accepted lower detection quality as a structural condition. That changes the risk model for IAM, PAM, and NHI telemetry because identity signals are only useful if they are reviewed in time. Security leaders should treat analyst capacity as part of control design, not as a back-office staffing issue.
Overtuned automation without local context creates detection debt. The more an SOC depends on out-of-the-box rules, the more false positives accumulate and the less trustworthy the queue becomes. That is a classic control failure because teams stop trusting the controls they paid for. Practitioners should understand that signal quality is an operational asset, not a static product feature.
Detection-response latency is the named concept this article sharpens: the time gap between suspicious activity and a human decision is now a measurable security exposure. In identity-heavy environments, that gap determines whether privilege misuse is contained or allowed to continue. Teams should measure latency by incident class, not just by overall response averages, because identity abuse is often fast and subtle.
The build-versus-buy debate is really a question of resilience under load. The article shows that some organisations will not achieve stable 24x7 coverage internally without disproportionate spend and burnout risk. That does not make in-house SOCs wrong, but it does mean leaders must re-evaluate which functions need direct ownership and which can be augmented or outsourced. The practitioner conclusion is to design for continuity, not pride.
IAM, PAM, and SOC teams need a shared operating picture. Identity events are among the highest-value indicators for threat hunting, yet they are also easy to bury in noisy operations. Mature governance requires that privileged access, service account activity, and anomalous authentication patterns flow into investigation workflows with clear ownership. The practitioner conclusion is to make identity telemetry a first-class SOC input, not a separate reporting stream.
What this signals
Detection-response latency is becoming a board-relevant risk metric because it captures whether the SOC can still convert telemetry into action before an incident expands. In identity-heavy environments, the practical challenge is not collecting more logs but shrinking the time between suspicious access and human review.
SOC leaders should expect more scrutiny over utilisation, backlog, and false-positive rates as organisations link operating cost to incident outcomes. That makes IAM and PAM telemetry more valuable, because identity anomalies are among the clearest indicators that the response model is falling behind.
For teams building programme plans, the real question is whether the SOC can still protect identity flows at peak load. If privileged access events, service account anomalies, and authentication outliers are not visible in the response queue, the programme is operating with a control gap rather than a staffing problem.
For practitioners
- Measure true analyst capacity Calculate average handling time, daily alert volume, and target utilisation before requesting headcount or tool budget. Use the numbers to determine whether the current model can sustain 24x7 coverage without pushing analysts above burnout thresholds.
- Tighten detection tuning before expanding tooling Review default rules, remove low-value alerts, and retune detections around your own environment and business processes. The goal is to reduce false positives so analysts can investigate identity abuse and genuine incidents with enough depth.
- Build identity telemetry into SOC triage Prioritise privileged account changes, service account misuse, and abnormal authentication events in the same workflow as other high-fidelity incidents. This shortens the path from signal to decision when attackers target credentials or access paths.
- Use hybrid coverage for repetitive triage work Keep internal ownership for threat hunting, escalation, and security architecture while offloading routine first-line review or after-hours queue handling where appropriate. This preserves expertise without forcing every coverage need onto the same team.
Key takeaways
- High alert volume and understaffing turn SOC efficiency into a security control issue, not just an HR problem.
- Identity events become harder to detect when analysts are overloaded, which weakens response to privilege misuse and suspicious access.
- Data-driven sizing, better tuning, and hybrid coverage are the practical levers that can restore sustainable operations.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
MITRE ATT&CK address the attack surface, NIST CSF 2.0, NIST SP 800-53 Rev 5 and CIS Controls v8 set the technical controls, and ISO/IEC 27001:2022 define the regulatory obligations.
| Framework | Control / Reference | Relevance |
|---|---|---|
| NIST CSF 2.0 | DE.CM-1 | Continuous monitoring is central to alert triage and SOC visibility. |
| NIST SP 800-53 Rev 5 | SI-4 | Security monitoring directly applies to alert overload and detection quality. |
| CIS Controls v8 | CIS-8 , Audit Log Management | Log management and review quality drive the SOC workload described here. |
| ISO/IEC 27001:2022 | A.8.16 | Monitoring activities are relevant to sustaining effective detection and response. |
| MITRE ATT&CK | TA0007 , Discovery; TA0006 , Credential Access; TA0040 , Impact | Alert overload enables attacker discovery, credential abuse, and delayed impact. |
Align log review workflows to CIS-8 and suppress low-value alerts that do not support response.
Key terms
- Alert Fatigue: Alert fatigue is the condition where a security team receives so many low-value alerts that important events become harder to notice. In monitoring programs, it usually signals poor rule tuning, weak prioritisation, or a mismatch between detection logic and operational reality.
- Analyst Gross Utilisation Rate: Analyst gross utilisation rate is the share of an analyst’s working time spent on active SOC work rather than idle or administrative time. When it stays too high for too long, it usually indicates an unsustainable workload, increasing the risk of burnout and degraded investigations.
- Detection-Response Latency: The elapsed time between identifying a security issue and executing a bounded, auditable fix. In data security programmes, long latency means exposure persists after discovery, which undermines the value of detection and weakens compliance evidence.
- Signal-to-Noise Ratio: The balance between meaningful security events and routine activity in detection tooling. A weak ratio makes analysts spend more time filtering alerts and less time identifying real attacks, which is why architecture quality strongly affects SOC effectiveness.
What's in the full article
Expel's full article covers the operational detail this post intentionally leaves for the source:
- A worked SOC cost calculator that translates alert volume, handling time, and utilisation into headcount requirements.
- Detailed salary and tooling cost ranges for basic, intermediate, and advanced SOC operating models.
- The specific warning signs that indicate a SOC is already beyond sustainable capacity.
- Practical build-versus-buy guidance for organisations deciding between internal, hybrid, and managed coverage.
👉 Expel's full article includes the capacity math, burnout indicators, and coverage model comparisons.
Deepen your knowledge
NHI Foundation Level course, the industry's only accredited NHI security programme, covers NHI governance, machine identity security, secrets management, and workload identity. It helps practitioners connect identity controls to the operational realities their programmes depend on.
Published by the NHIMG editorial team on August 20, 2026.
NHI Mgmt Group — the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org