Join our Newsletter — 33% off our NHI Course
Home› FAQ› Threats, Abuse & Incident Response› What are the signs that behavioural baselines are…
Threats, Abuse & Incident Response

What are the signs that behavioural baselines are not working properly?

← Back to all FAQ
By NHI Mgmt Group Editorial Team Updated October 11, 2026 Domain: Threats, Abuse & Incident Response

Common warning signs are baselines with too little history, scores that treat every entity the same, and detections that cannot be explained from the underlying data. If the model cannot show what was learned, or if it still pages on normal behaviour for one entity while missing the same pattern on another, the baseline design is too coarse.

What a failing behavioural baseline usually looks like in practice

When behavioural baselines are working, they separate routine from unusual activity in a way operators can trust. When they are failing, the warning signs usually show up as unstable alerting, poor explanation quality, or an inability to distinguish one entity’s normal pattern from another’s. The problem is rarely “no model”; it is usually a model that learned the wrong shape of normal.

A practical test is whether the baseline still behaves sensibly when the same action is seen across different users, hosts, workloads, or sessions. If the signal flips from noisy to silent without a meaningful change in context, the baseline is likely over-smoothed, over-generalised, or trained on too little stable history.

Another sign is that the output is technically scored but not operationally interpretable. If analysts cannot connect a detection back to the features, time window, or prior behaviour that drove it, the baseline may be fitting correlations without producing a defensible security judgement. That is a design failure, not just a tuning issue.

Why history, context, and entity separation matter

Behavioural baselines depend on enough clean history to distinguish recurring routine from meaningful deviation. Short observation windows, seasonal shifts, onboarding periods, and shared accounts can all distort what “normal” looks like. A baseline built on weak history will often appear precise in a dashboard while failing in live operations.

The other common failure is treating all entities as though they belong to one population. A baseline for a developer workstation, a production server, and a batch process should not behave identically, even if the same control plane is watching them. When the model ignores entity role, location, or purpose, it produces false positives for the active populations and blind spots for the quieter ones.

That is why a good baseline must preserve meaningful separation in the data, while still borrowing enough context to avoid learning one-off noise. In security terms, the goal is not statistical elegance, it is durable discrimination.

How to tell whether the baseline is actually learning the right behaviour

The strongest indicator of a healthy baseline is consistency between the learned pattern and the underlying evidence. You should be able to explain why a given event was scored as normal or abnormal, and that explanation should remain stable when the same event is replayed against the same entity. If the model cannot show what it learned, the team cannot validate whether it learned the right thing.

Model quality also shows up in comparative behaviour. If one entity triggers on expected admin work while another entity with a similar operating profile is ignored, the baseline is not learning risk, it is learning uneven coverage. That often means the training data is incomplete, the entity segmentation is too coarse, or the feature set does not capture the right operational signals.

For practitioners who want a hard reference point on baseline quality and operational hardening, CIS Benchmarks are a useful companion because they show how control expectations depend on the system and platform being measured, not on a single universal standard. CIS Benchmarks help anchor the idea that “normal” must be defined per technology and context.

Risk and Threat Considerations

Broken baselines create two kinds of exposure: alert fatigue when normal activity is repeatedly misread, and silent failure when genuinely odd activity blends into an over-broad model. In adversarial environments, attackers benefit from either condition because they can hide inside noisy detections or exploit blind spots created by poor entity separation.

Failure mechanism: The baseline learns from insufficient or mixed history, or it collapses distinct entities into one behaviour profile, so the system cannot distinguish stable routine from meaningful deviation.

Impact: Teams waste time on false alerts, trust the model less, and may miss the very behavioural change they intended to catch, especially when the same action is treated differently across entities.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

MITRE ATT&CK addresses the attack and risk surface, while CIS Controls v8, NIST CSF 2.0 and NIST SP 800-53 Rev 5 set the governance and control requirements practitioners need to meet.

FrameworkControl / ReferenceRelevance
CIS Controls v8CIS-5 — Account ManagementBaseline quality depends on distinct entity roles and histories.
Recommendation — Segment monitoring by account or entity role before tuning anomaly thresholds.
NIST CSF 2.0DE.AE-02 — Anomalous activity is detected and analysedBehavioural baselines are anomaly detection mechanisms.
Recommendation — Validate that anomaly logic distinguishes normal from suspicious entity behaviour.
NIST SP 800-53 Rev 5AU-6 — Audit Record Review, Analysis, and ReportingBaseline failures surface when detections cannot be explained from the data.
Recommendation — Require reviewable evidence for why a behaviour was scored as anomalous.
MITRE ATT&CKT1036 — MasqueradingAttackers exploit weak baselines by blending into expected behaviour patterns.
Recommendation — Map suspicious routine-like activity to masquerading and hunt for blending tactics.

Practitioner Guidance

What to verify: Check whether the model can justify each alert using the entity’s own history, not just aggregate population behaviour. If the explanation changes wildly between reruns or cannot be traced to specific learned features, treat the baseline as untrusted.

Common mistake: Teams often optimise for broader coverage before proving that the baseline is stable per entity class. That usually increases noise faster than it increases detection value.

What good looks like: A sound baseline produces repeatable scoring, different thresholds or profiles where operational roles differ, and a clear reason for why one entity is abnormal while another similar one is not.

Practitioner takeaway: If the baseline cannot explain itself and cannot distinguish between entities with different normal behaviour, the next step is not more tuning, it is redesigning the training scope and entity model.

Free weekly newsletter

Subscribe to the NHI & AI Identity Journal

The latest on NHI and Agentic AI security – articles, research, breaches, news and events every week.

Bonus 33% off our NHI Course when you subscribe.

NHIMG Editorial Note
Reviewed and updated by the NHIMG editorial team on October 11, 2026.
NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org