The clearest signs are inconsistent login behaviour, health checks that pass on one node but not another, cluster size mismatches, and intermittent errors after traffic shifts. Those symptoms usually indicate that database replication, session storage, or configuration parity is incomplete rather than that the application itself is unstable.
How to tell the cluster is masking a replication or session problem
A clustered Passbolt deployment is usually only as available as the state it shares between nodes. If users can authenticate on one node and then fail or see stale behaviour after a traffic shift, the problem is often not the application binary, it is the synchronisation layer. The first thing to distinguish is whether the failure follows the node, the session, or the backend state.
That distinction matters because a “working” node can still sit behind an incomplete cluster if shared storage, database replication, or configuration propagation is lagging. When the symptom changes with node selection or load balancer routing, you are looking at a consistency gap, not a random service outage.
Healthy looking health checks can also mislead. A node may report as up while another node has divergent data, an expired session store, or a partially applied configuration update. In practice, high availability means the NIST Cybersecurity Framework 2.0 style expectation of resilient service delivery only holds when the cluster state is genuinely interchangeable across nodes.
Which symptoms point to incomplete cluster parity
The strongest indicators are inconsistent login behaviour, health checks that disagree across nodes, and intermittent errors that appear only after failover or traffic redistribution. Those symptoms suggest the cluster is not presenting one coherent service state to users.
Cluster size mismatches are another practical warning sign. If one node is missing from rotation, is not joining cleanly, or is showing a different version or configuration, the cluster may look redundant on paper while still behaving like a partial single point of failure in production.
For a cloud-hosted deployment, the control question is whether the deployment is matching the intended operational pattern for shared state, access, and isolation. A CSA Cloud Controls Matrix lens is useful here because it forces you to inspect whether access, infrastructure, and data handling are actually consistent across the environment rather than assumed to be.
What the failure pattern usually means in practice
When the symptoms are intermittent and node-specific, the likely issue is that one of three dependencies is not truly clustered: the database, the session layer, or the configuration/state distribution mechanism. Any one of those can make a deployment appear redundant while still breaking under node changes.
Session problems are especially easy to miss because they often show up only after a redirect or failover. If authentication succeeds but subsequent requests lose context, the deployment may not be preserving session affinity, or the session backing store may not be shared and durable enough for failover.
Configuration drift can create a similar illusion. If one node has a different secret, cookie setting, base URL, or worker setting, the cluster may pass superficial checks but fail under real user flow. That is why the operational model behind NIST SP 800-53 Rev 5 Security and Privacy Controls is relevant: consistency, access control, and configuration integrity are core availability enablers, not separate concerns.
Risk and Threat Considerations
A cluster that is “available” only on some nodes can create hidden outage risk and, in a security-sensitive application, hidden trust risk. Users may retry, refresh, or reauthenticate into a degraded state that looks normal until a failover exposes missing shared state or broken parity.
Failure mechanism: The deployment treats node uptime as proof of high availability, but the shared dependencies that make nodes interchangeable are not fully consistent. Database lag, session fragmentation, or configuration drift then surface only when traffic moves.
Impact: Failover can turn a seemingly healthy deployment into an intermittent outage, with failed logins, broken workflows, and inconsistent access behaviour that is hard to diagnose quickly.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
NIST CSF 2.0, CSA Cloud Controls Matrix and NIST SP 800-53 Rev 5 set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| NIST CSF 2.0 | RC.RP-01 — Recovery Plan Execution | Cluster failover symptoms map to recovery behaviour and service restoration. |
| Recommendation — Test failover paths to confirm the cluster restores service without state loss. | ||
| CSA Cloud Controls Matrix | IAM — Identity & Access Management | Passbolt availability depends on consistent access, session, and node behaviour across the cluster. |
| Recommendation — Verify access and session controls remain consistent across all nodes. | ||
| NIST SP 800-53 Rev 5 | SC-5 — Denial of Service Protection | Intermittent errors after traffic shifts expose resilience and service continuity weaknesses. |
| Recommendation — Validate that load shifts do not create avoidable service interruption. | ||
Practitioner Guidance
What to verify: Test the exact user journey across node changes, not just service liveness. A deployment is not truly high available until authentication, session continuity, and stateful requests survive a node shift without behavioural changes.
What to prioritise: Check the shared dependencies first, especially database replication status, session store durability, and configuration parity. If those are not aligned, application troubleshooting will waste time and obscure the real fault domain.
Practitioner takeaway: Treat “passes health check” as a weak signal; the stronger test is whether the cluster preserves the same user-visible state after loss of any single node.
Related resources from NHI Mgmt Group
- What are the main signs that IoT identity and connectivity controls are not keeping pace with deployment growth?
- What are the signs that an AI system is not ready for high-risk deployment under the EU AI Act?
- What are the signs that certificate deployment is failing in a clustered environment?
- What are the main signs that a digital identity process is not suitable for high-pressure service environments?