By NHI Mgmt Group Editorial TeamBased on Pomerium: “7 Things to Know About Kubernetes Health Checks” (September 9, 2025)

TL;DR: Kubernetes health checks use startup, readiness, and liveness probes to manage containers, but Pomerium argues that eventual consistency, stateful dependencies, and split-mode deployments make those signals harder to trust in practice, according to Pomerium. The governance lesson is that health checks are only as reliable as the assumptions behind them, especially when access, proxying, and observability overlap.


At a glance

What this is: This article explains why Kubernetes health checks become unreliable when stateless probe logic meets stateful services, eventual consistency, and multi-subsystem deployments.

Why it matters: It matters because IAM, NHI, and platform teams increasingly depend on health signals to govern access, routing, and recovery decisions, and weak probes can hide real operational failure.


Context

Kubernetes health checks are not just a reliability feature. They are a control signal that tells the platform when to start, stop, or restart workloads, which means bad probe design can create false confidence about service state.

The governance gap is that probe logic often assumes applications behave like simple stateless units, while real services depend on caches, upstream systems, and synchronised internal components. Once those dependencies diverge, readiness and liveness stop describing the full operational picture.

That tension is especially visible in platform components that sit between users and back-end services, where access decisions, proxy state, and observability all overlap.


Key questions

Q: Where do Kubernetes health checks fail in stateful environments?

A: They fail when readiness and liveness are expected to represent complex internal state with a single binary signal. Stateful services can be running while still not safe to receive traffic, or briefly unhealthy during normal convergence. Teams need subsystem-aware checks and observability that explain what the probe cannot capture.

Q: Why do eventual consistency and probe timing create operational risk?

A: Because the platform may react before the application has finished converging across caches, configuration, or upstream dependencies. That can trigger premature restarts or traffic shifts, even when the underlying service would have stabilised on its own. The risk is misclassification, not just noise.

Q: How should teams diagnose a pod that is repeatedly marked unhealthy?

A: Start by checking whether the failing signal is local, upstream, or caused by a transitional state such as rollout convergence or cache invalidation. Then compare probe results with traces, logs, and dependency health so you can distinguish a real service fault from a timing problem.

Q: Should readiness checks be the only control for routing traffic?

A: No. Readiness is useful, but it should not be the sole routing decision when services have multiple subsystems or asynchronous dependencies. A better model combines readiness with richer operational evidence so traffic moves only when the service is actually safe to serve.


Technical breakdown

Why Kubernetes probes struggle with stateful services

Kubernetes uses three core probes. Startup checks whether an application has finished initialisation, readiness determines whether it should receive traffic, and liveness determines whether it should be restarted. These probes work well when state is local and predictable, but they become blunt instruments when the service depends on upstream systems, caches, or asynchronous configuration. In those cases, a pod can be alive but not ready, or appear ready while critical subsystems are still converging. The problem is not the probe API itself, but the assumption that health can be reduced to a binary signal across complex distributed states.

Practical implication: model health at subsystem level, not only at container level, so readiness reflects the state that actually matters.

How eventual consistency distorts readiness and liveness

Eventually consistent systems do not guarantee that every component sees the same state at the same moment. Kubernetes continuously polls probe endpoints and changes routing or restarts workloads after thresholds are crossed, which makes timing part of the control plane. If cache invalidation, config propagation, or dependency synchronisation lags behind the probe window, the platform can react to transitional failure as if it were permanent. That creates oscillation: traffic shifts, pods restart, and operators lose confidence in the signal. In practice, a health check becomes a policy decision made on incomplete state rather than a faithful measurement of service viability.

Practical implication: set probe thresholds around propagation delay and recovery behaviour, not around idealised application timing.

Why split-mode deployments need richer observability than probes alone

Split-mode deployments introduce separate operational paths for authentication, authorisation, proxying, or other service functions, so one subsystem can fail while the others still appear healthy. Pomerium’s example shows why operators need more than a green or red probe. Metrics, logs, and traces provide the context that probes cannot, especially when failures originate upstream or only affect one service path. OpenTelemetry-style tracing helps distinguish between a platform problem and an external dependency problem. The architectural lesson is that health checks are a gate, but observability is the diagnostic layer that explains what the gate is actually seeing.

Practical implication: pair probes with traces and error context so operators can separate real service failure from upstream dependency failure.


NHI Mgmt Group analysis

Probe-based governance breaks when service state is no longer singular. Kubernetes health checks assume a workload can be classified cleanly as starting, ready, or unhealthy. That assumption weakens when authentication, proxying, caching, and configuration synchronisation each move on different timelines. The practitioner takeaway is that reliability controls must match the actual state model of the service, not the convenience of the probe interface.

Eventual consistency creates an identity-style trust problem inside infrastructure operations. Operators end up trusting a signal that may only describe one subsystem, one cache, or one moment in the propagation cycle. That is a governance problem as much as an engineering one, because access, routing, and recovery decisions depend on whether the signal truly represents the service boundary. The implication is to treat health as an evidence chain, not a single status code.

Split-mode architectures amplify observability debt. When different parts of a platform can fail independently, a binary health answer obscures where responsibility sits and why traffic changed. Clear tracing, structured logs, and per-subsystem checks become the practical minimum for trustworthy operations. The practitioner conclusion is simple: if the architecture is modular, the health model must also be modular.

Kubernetes health checks are only as trustworthy as the assumptions behind them. The article sharpens a useful concept here: health signal credibility. If probes are calibrated for simple apps, they will mislead operators in complex, stateful environments. Teams should therefore align probe design, observability, and deployment topology before treating health as an automatic source of truth.

From our research library:

What this signals

Health checks should be treated as one input into operational decision-making, not as a complete description of service state. In distributed environments, the more a platform depends on asynchronous propagation, the more likely binary probe logic is to hide useful detail.

The practical lesson for identity and platform teams is that trust in a service signal must be earned through correlation. When readiness, tracing, and dependency visibility all agree, operators can act with confidence; when they diverge, the signal is telling you the model is incomplete.


For practitioners

  • Map each critical subsystem to its own health signal Define separate checks for authentication, authorisation, proxying, cache state, and dependency reachability so one subsystem does not mask another.
  • Tune probe thresholds around real propagation delays Set startup, readiness, and liveness thresholds using measured synchronisation times, restart behaviour, and upstream recovery windows rather than default values.
  • Use traces to explain unhealthy states Add distributed tracing and structured logs so operators can see whether failure is local to the service or inherited from an upstream dependency.
  • Separate transient convergence from genuine failure Treat cache invalidation, rollout convergence, and config propagation as expected transitional states, and avoid mapping them directly to restart logic unless the evidence supports it.

Key takeaways

  • Kubernetes health checks become unreliable when a service’s real state is spread across caches, upstream dependencies, and independently failing subsystems.
  • Binary probes can cause premature restarts or traffic shifts if they do not account for propagation delay and convergence behaviour.
  • Teams should pair probe design with richer observability so health decisions reflect actual service viability rather than a simplified status code.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

OWASP API Security Top 10 addresses the attack and risk surface, while NIST SP 800-190, NIST CSF 2.0 and CIS Controls v8 set the governance and control requirements practitioners need to meet.

FrameworkControl / ReferenceRelevance
NIST SP 800-190Container SecurityKubernetes probe design and workload observability sit inside container security operations.
Recommendation — Review container health and observability controls so probes reflect the real state of running workloads.
NIST CSF 2.0PR.PS-04 — Platform SecurityHealth checks are part of platform assurance for containerised services.
Recommendation — Align platform monitoring with service state so routing and recovery decisions are based on trusted signals.
CIS Controls v8CIS-8 — Audit Log ManagementTracing and structured logs are needed to explain probe failures and dependency issues.
Recommendation — Use logging and trace correlation to distinguish transient convergence from genuine service failure.
OWASP API Security Top 10API8 — Security MisconfigurationMisconfigured probe logic and service endpoints can produce unsafe operational decisions.
Recommendation — Validate health endpoints and routing logic so configuration errors do not masquerade as availability.

Key terms

  • Kubernetes Cluster Health Monitoring: Kubernetes cluster health monitoring is the practice of checking whether the platform itself is functioning, not just whether workloads are deployed. It focuses on control-plane services, networking, storage behavior, and failure recovery so operators can detect degraded clusters before applications fail in ways that simple status checks miss.
  • Eventual Consistency: A system property where a change is accepted before every part of the platform reflects it. In cloud identity, that matters because a revoked permission may still be usable for a short period, creating a window in which an attacker or automated tool can act before enforcement converges.
  • Readiness probe: A readiness probe checks whether a service can safely receive traffic, not merely whether it is running. In distributed systems, it should reflect initialization, dependency sync, and control-path integrity so orchestration does not route requests into a partially prepared instance.
  • Liveness probe: A liveness probe checks whether a process is still functioning and should be restarted if it is not. It is useful for crash detection, but it does not prove the service is ready to serve requests or that its internal state is consistent.

Deepen your knowledge

NHI governance, agentic AI identity, and machine identity lifecycle are core topics in our NHI Foundation Level course, the industry's only accredited NHI security programme. If you are responsible for identity security strategy or NHI governance in your organisation, it is worth exploring.
NHIMG Editorial Note
Published by the NHIMG editorial team on June 11, 2026.
Updated on October 10, 2026.
NHI Mgmt Group, the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org