TL;DR: Stack-aware health checks lifted client query success from near zero to 99.9% during scaling, restarts, and rollouts by checking readiness across Kubernetes, Docker, and systemd instead of relying on basic process liveness, according to Pomerium. The deeper lesson is that availability controls fail when they measure uptime but not operational readiness.
At a glance
What this is: This is a Pomerium engineering analysis of why basic health checks failed to reflect proxy readiness and how stack-aware checks improved real service availability.
Why it matters: IAM and infrastructure teams need this because zero trust controls depend on readiness signals that match actual request path behavior, not just whether a process is running.
Context
A health check only helps if it measures the state that actually determines whether a zero trust proxy can serve traffic. In Pomerium’s case, the control gap was not process uptime but readiness across multiple layers that had to agree before requests could succeed.
The article frames this as a lifecycle problem inside a stateful proxy: Kubernetes, Docker, and systemd all surface health differently, but only some signals are strong enough to support self-healing behavior. That makes the problem relevant to IAM teams running zero trust gateways, where the control plane can look healthy while request paths are still unstable.
Key questions
Q: What breaks when a zero trust proxy is marked healthy before it is ready to serve traffic?
A: The proxy can enter rotation while authentication, authorization, or backend state is still incomplete, so users see failed requests even though orchestration reports success. The failure is a mismatch between lifecycle state and access-path state. Teams should define readiness around the exact conditions needed to make a correct decision about every request.
Q: Why do readiness checks matter more than liveness checks for proxy availability?
A: Liveness only shows that a process is still running, which is not enough for a stateful access gateway. Readiness confirms that the service can safely accept requests after its internal dependencies, configuration, and policy state are coherent. Without that distinction, scaling and rollout events can create avoidable traffic loss.
Q: How can security teams tell whether a proxy health model is too shallow?
A: If the health signal only tracks process survival, container status, or a single endpoint, it is probably too shallow for a zero trust gateway. A better model captures the conditions that govern request admission, state synchronization, and clean termination. That tells you whether health is operationally meaningful or just cosmetically green.
Q: What should teams review before using health checks to automate self-healing?
A: They should review whether the check measures the exact state needed to admit traffic, not a proxy for it. Self-healing is only safe when the signal maps to the service’s real operating boundary. Otherwise the automation can restart or promote instances that are still unsuitable for production use.
Technical breakdown
Why liveness checks are not readiness checks
Liveness tells an orchestrator that a process is still running. Readiness tells it whether the process can safely receive traffic. In a zero trust proxy, those are not equivalent because authentication, authorization, proxying, and storage all have to align before the first request can succeed. Pomerium’s article shows that a container can be alive while initialization, state sync, or dependency setup is still incomplete. Kubernetes distinguishes startup, readiness, and liveness probes for this reason: they answer different operational questions at different points in the lifecycle.
Practical implication: treat liveness as a restart signal, not as proof that a zero trust service can carry production traffic.
How stack-aware health checks model proxy state
A stack-aware health check combines signals from the layers that actually determine service behavior. In Pomerium’s design, that includes Kubernetes probes, Docker health status, and systemd notifications, each reflecting a different operating context. The key idea is that health must be modeled against the service’s own state machine, not a generic process heartbeat. That is especially important when configuration is hot-reloadable and state is synchronised across multiple backend services, because the proxy can be running while its internal view of policy and routing is still catching up.
Practical implication: define health around the service state that determines access decisions, then expose that state through the orchestrator layer you actually run.
Why eventually consistent systems need stricter readiness thresholds
Eventually consistent systems can report intermediate states that are technically valid but operationally unsafe. Pomerium addresses that by using failure and success thresholds so a transient state does not immediately flip the service in or out of rotation. That matters when configuration records are applied in ordered sequence and the proxy must keep auth, authorization, and upstream routing aligned during scale events. The design trade-off is clear: coarse checks are easier to implement, but they hide the exact failure mode that causes traffic loss during rollout or restart.
Practical implication: require readiness thresholds that prevent partially synchronized instances from entering the active pool too early.
NHI Mgmt Group analysis
Readiness is a governance control, not just an uptime check: zero trust proxies expose a control assumption that many teams miss, which is that a service can be considered available once the process exists. That assumption breaks when authentication, authorization, and state synchronization are not yet aligned. The practitioner takeaway is to govern service entry into traffic based on functional readiness, not container presence.
Stack-aware health checks reduce false confidence in availability reporting: the article shows that process-level signals can misstate the readiness of a proxy that sits in the access path. When a gateway is stateful, the health model has to reflect the actual request chain or it will certify the wrong thing. This is a service governance issue as much as an operations issue, and teams should treat it that way.
Zero trust architectures are only as strong as their state transitions: if an ingress proxy can be scaled, restarted, or reloaded before policy and session state are coherent, the access boundary becomes unstable. The real failure is not absence of a check, but mismatch between the check and the stage of the service lifecycle it is meant to control. Practitioners should align health semantics with the state transitions that gate traffic.
Operational readiness becomes the decisive control variable in stateful proxies: the article’s strongest insight is that availability logic must account for initialization, synchronization, and termination, not only runtime success. That shifts the question from whether a service is up to whether it is fit to arbitrate access at that moment. Teams managing zero trust gateways should measure the readiness boundary itself, not the proxy binary.
Pomerium’s design illustrates a broader identity operations pattern: access infrastructure fails when orchestration tools assume that a running component is also a trustworthy one. The article shows why readiness signals need to be specific to the service’s own internal state, especially when configuration and policy are hot-reloaded. The practitioner implication is to make readiness semantics part of identity platform design, not an afterthought.
What this signals
Operational readiness is the boundary that matters: access infrastructure should not be judged healthy until the service can actually authorize and route requests. That distinction matters most in stateful zero trust proxies, where lifecycle transitions can briefly expose partially initialized state to the orchestration layer.
Health semantics need to reflect the service state machine: Kubernetes-style probes, Docker health status, and systemd notifications all expose different levels of truth. Teams should decide which state transition each signal is supposed to control, then avoid using a coarse process check as a surrogate for traffic admission.
Zero trust proxy governance should include readiness semantics: when policy, session, and backend state are hot-reloaded, the service can appear up while still being unfit to broker access. Practitioners should align monitoring, rollout procedures, and traffic gating with the exact point at which the proxy becomes trustworthy for live requests.
For practitioners
- Define readiness by request-path eligibility Base health on whether the proxy can actually authenticate, authorize, and route traffic successfully, not whether the process is merely running.
- Separate startup, readiness, and liveness Use different checks for initialization, traffic admission, and crash detection so the orchestrator does not conflate recovery with availability.
- Gate traffic on synchronized state Prevent new instances from joining the active pool until policy, configuration, and backend state have reached a coherent operating point.
- Treat hot reloads as health-sensitive events Validate that configuration reloads do not expose partially applied state to the access path before the proxy resumes full service.
Key takeaways
- Basic health checks can report a proxy as alive while its access path is still not ready, which creates avoidable request failures during scale events and rollouts.
- The central issue is operational readiness, not simple process uptime, and the article shows why state-aware checks produce a more accurate availability signal.
- Teams running zero trust gateways should tie traffic admission to synchronized service state so orchestration does not promote instances too early.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Non-Human Identity Top 10 addresses the attack and risk surface, while NIST CSF 2.0, NIST Zero Trust (SP 800-207) and CIS Controls v8 set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP Non-Human Identity Top 10 | NHI-06 — Insecure Cloud Deployment Configurations | The article is about misaligned health semantics in cloud orchestration and proxy lifecycle handling. |
| Recommendation — Align deployment health signals with the service state required before traffic admission. | ||
| NIST CSF 2.0 | PR.AA-05 — Access Permissions, Entitlements and Authorizations | Readiness controls determine when access decisions may safely be enforced for live traffic. |
| Recommendation — Tie access admission to verified service readiness before placing instances in rotation. | ||
| NIST Zero Trust (SP 800-207) | Zero Trust Architecture — Zero Trust Architecture | The article directly examines operational readiness inside a zero trust proxy. |
| Recommendation — Map proxy readiness checks to the trust boundary they are intended to protect. | ||
| CIS Controls v8 | CIS-5 — Account Management | Lifecycle state and service admission are both governance problems for access-bearing systems. |
| Recommendation — Review whether account and service state changes can occur before the system is ready. | ||
Key terms
- Readiness Checklist: A readiness checklist is a structured control tool used to verify that required evidence, approvals, and processes are in place before an AML audit. It helps teams confirm coverage across policies, monitoring, escalation, and remediation. A useful checklist is operational, current, and tied to real ownership.
- Active Liveness Check: An active liveness check requires the user to complete a prompt during verification, such as blinking, smiling, or turning the head. The system uses the response to confirm presence and detect replay or spoofing attempts. It is stronger against fraud, but it adds friction and depends on user participation.
- Zero Trust Browser: A Zero Trust Browser is a browser designed to reduce trust in the endpoint and the web session itself. It applies identity, policy, inspection, and isolation controls to browser activity, so access to web apps and data is continuously evaluated rather than assumed safe after login.
- Eventual Consistency: A system property where a change is accepted before every part of the platform reflects it. In cloud identity, that matters because a revoked permission may still be usable for a short period, creating a window in which an attacker or automated tool can act before enforcement converges.
Deepen your knowledge
NHI governance, agentic AI identity, and machine identity lifecycle are core topics in our NHI Foundation Level course, the industry's only accredited NHI security programme. If you are building or maturing an IAM programme, it is worth exploring.
Published by the NHIMG editorial team on June 11, 2026.
Updated on October 10, 2026.
NHI Mgmt Group, the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org