Alerting on bad things watches for thresholds, errors, and failures. Alerting on good things watches for normal activity that suddenly drops, such as logins, page loads, or successful actions. The second approach catches silent failures that may not raise errors, while the first is better at identifying obvious breaks and resource exhaustion.
Alerting on Failures Versus Alerting on Silence
Alerting on bad things is event-driven: you fire when an error, threshold breach, or explicit failure appears. Alerting on good things that stop happening is absence-driven: you watch for a normal signal and alert when it disappears or drops below expectation. The difference matters because some outages fail loudly, while others silently break the healthy activity you expected to keep seeing.
The first pattern is easiest to operationalise for systems that emit clear error states, saturation signals, or policy violations. The second pattern is more useful where “no news” is itself a symptom, such as a scheduler that stops running, a login flow that quietly stalls, or a message pipeline that no longer produces successful completions. The detection strategy should match the failure mode, not just the telemetry you happen to have.
A useful way to think about the distinction is that bad-thing alerting detects explicit harm, while good-thing-stopped alerting detects loss of liveness. The latter often finds problems earlier when the system has not yet reached a hard failure state. It also protects against blind spots where a component can break in a way that prevents the expected success signal from ever being emitted.
Why Silent Failures Need Different Thresholds
Silence-based alerting usually needs a baseline, a time window, and a confidence threshold. You are not measuring whether something is “too high,” but whether an expected pattern has become unusually sparse or absent. That means normal diurnal variation, maintenance windows, and low-volume edge cases must be accounted for, otherwise the alert stream becomes noisy and teams stop trusting it.
The design choice is whether the signal is a heartbeat, a success counter, or a business event. A heartbeat tells you the component is still alive. A success counter tells you the workflow is still completing. A business event tells you the system is still delivering value. Each has different sensitivity, and the right choice depends on what failure would be most dangerous to miss.
Good-thing alerting is especially strong when failure can mask itself behind retries, queues, or fallback behaviour. A system may continue processing without throwing obvious errors, but the rate of successful logins, transactions, page loads, or downstream acknowledgements can still collapse. That is why absence-based monitoring often catches partial outages, integration failures, and routing defects sooner than error-only monitoring.
Choosing the Right Signal for Operational Reality
Use bad-thing alerts when the failure is explicit, bounded, and meaningful in its own right, such as a service crash, elevated error rate, or resource exhaustion. Use good-thing-stopped alerts when continuity is the risk, such as a missing heartbeat, a sudden drop in successful actions, or a pipeline that should produce regular activity. In practice, mature monitoring uses both, because they observe different sides of the same operational problem.
The most effective setups anchor the silence alert to something the business actually expects to happen. A generic “no traffic” alert can be misleading, but “no successful logins from a critical identity provider for 10 minutes” or “no completed jobs from a scheduled worker since the last expected run” is actionable. The narrower and more meaningful the expected behaviour, the better the alert quality.
For practitioners, the key question is not which approach is better in the abstract, but which one detects the failure mode before users or downstream systems feel it. That is why many teams pair threshold alerts with liveness and success-signal alerts, then tune them separately. One protects against obvious breakage, the other against quiet degradation.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
NIST CSF 2.0, NIST SP 800-53 Rev 5 and CIS Controls v8 set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| NIST CSF 2.0 | DE.CM-01 — Networks and systems monitored to detect potential cybersecurity events | Alerting on missing or abnormal signals is continuous monitoring. |
| Recommendation — Monitor expected signals and alert on sustained departures from baseline. | ||
| NIST SP 800-53 Rev 5 | AU-6 — Audit Review, Analysis, and Reporting | Alerting depends on analyzing logs and events for loss of expected activity. |
| SI-4 — System Monitoring | Both failure alerts and silence alerts rely on system monitoring for detectable conditions. | |
| Recommendation — Review telemetry for missing success events and abnormal gaps. Define monitors for explicit faults and for missing expected activity. | ||
| CIS Controls v8 | CIS-8 — Audit Log Management | Good-thing-stopped alerts are typically built from log and event monitoring. |
| Recommendation — Collect and alert on missing or reduced log and event activity. | ||
Practitioner Guidance
What to prioritise: Treat absence alerts as first-class signals for critical workflows, not as an afterthought. If the success of a process matters more than the existence of errors, monitor the success signal directly.
What to verify: Confirm that the alert is tied to an expected cadence, not raw volume alone. A useful silence alert should answer, “What normal thing should have happened by now?”
Decision rule: If the failure can occur without generating an error, choose a good-thing-stopped alert. If the failure reliably produces explicit faults, error-based alerting is usually the cleaner primary signal.
Practitioner takeaway: The strongest monitoring mixes both perspectives, because systems often fail either by breaking loudly or by becoming quietly inactive, and those are not the same operational problem.
Related resources from NHI Mgmt Group
- What is the difference between good bots and bad bots?
- What is the difference between a good MCP tool definition and a bad one?
- What is the difference between access control and PCI alerting in SharePoint?
- What is the difference between a good benchmark and a useful benchmark for AI security scanners?