Teams should move from hidden session state to explicit state passed in requests. A server can mint an opaque handle, such as a workflowId or cursor, and return it from the first tool call. The client then supplies that handle on later calls. This makes workflow context visible, survives stateless routing, and avoids brittle dependence on connection affinity or a shared session store.
Redesign MCP workflows around explicit request state
Once protocol-level sessions are removed, the workflow has to stop depending on hidden continuity and start carrying its own state. The practical change is simple: treat each interaction as stateless transport plus an explicit workflow token, so the server can recover context without assuming the same connection, pod, or session store will still be present later.
This is a design shift, not just an implementation detail. If the workflow state is not visible in the request, the system becomes fragile under retries, horizontal scaling, failover, and client restarts. If the state is explicit, the protocol can survive routing changes while still giving the server enough context to continue the task correctly.
How to use handles, cursors, and workflow IDs safely
The clean pattern is for the server to mint an opaque handle, such as a workflowId or cursor, and return it from the first tool call. The client then supplies that handle on later calls, so the server can resume the right workflow without exposing internal implementation details. That keeps the protocol surface small and avoids making the client interpret server-side state formats.
An opaque handle should identify a server-owned workflow state, not become a shortcut for broad authority. The handle should be scoped to the specific workflow, validated on every use, and rejected when it is stale, malformed, or presented in the wrong context. If the handle can be replayed across unrelated tasks, the redesign has only replaced session state with a different kind of hidden coupling.
This approach also works well when workflows need pagination, incremental retrieval, or multi-step tool execution. A cursor or workflow ID lets the server continue from the last known point while still preserving the ability to rebalance traffic or recover from a failed node.
What breaks if you keep relying on session affinity
Connection affinity and shared session stores are the two usual failure points after sessions disappear. Affinity makes the design brittle because the next request may land on a different instance. A shared store can restore continuity, but it introduces another operational dependency, more synchronization cost, and a larger blast radius if the store is slow or unavailable.
The better mental model is that the workflow state belongs to the request contract, while the transport only moves bytes. That lets teams reason about correctness at the protocol boundary instead of relying on infrastructure behavior to keep the conversation alive. It also makes observability easier, because the workflow state can be logged, traced, and audited explicitly.
Risk and Threat Considerations
Explicit state reduces fragility, but it also creates a new object that must be protected like any other bearer of workflow context. If a workflow handle can be guessed, replayed, or reused outside its intended scope, a client may resume or influence work it should not control.
Failure mechanism: Long-lived or poorly validated handles can be replayed after a transport change, stolen from logs or client traces, or accepted across unrelated workflows. That creates confusion between continuity and authority, especially when the handle is treated as proof that the caller may continue a task.
Impact: The result can be unauthorized continuation, cross-workflow data exposure, incorrect task execution, or a broader trust failure if multiple services interpret the same handle differently. The safer design is to make the handle narrow, short-lived where possible, and bound to the exact workflow and caller context.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Agentic AI Top 10 and OWASP API Security Top 10 address the attack and risk surface, while NIST SP 800-53 Rev 5 sets the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP Agentic AI Top 10 | ASI02 — Tool Misuse | MCP workflows use tool calls and delegated steps that must be bounded per request. |
| ASI03 — Identity & Privilege Abuse | Workflow handles can become authority-bearing if reused across steps or contexts. | |
| Recommendation — Bind each resumed workflow step to explicit caller intent and validate tool use on every request. Scope workflow handles narrowly and reject reuse outside the original caller and task. | ||
| OWASP API Security Top 10 | API4 — Unrestricted Resource Consumption | Stateless workflow redesign must still prevent cursors or handles from enabling repeated or amplified calls. |
| Recommendation — Rate-limit resumed workflow paths and cap stateful retries to prevent abuse. | ||
| NIST SP 800-53 Rev 5 | IA-5 — Authenticator Management | Workflow handles and related secret material need lifecycle controls, expiry, and revocation. |
| AC-6 — Least Privilege | Opaque handles should not grant broad access beyond the specific workflow state they resume. | |
| Recommendation — Expire and rotate workflow handles or related tokens when the workflow ends. Limit each handle to the minimum workflow actions needed for continuation. | ||
Practitioner Guidance
What to verify: Confirm that the server can reconstruct all required workflow context from the request alone, without relying on sticky routing or a hidden session cache. If a later step needs more than the returned handle, add that state to the explicit contract instead of assuming infrastructure will preserve it.
Implementation sequence: Start by identifying the minimum state needed to resume a workflow, then define the handle format, expiry, and validation rules, and only then remove the session dependency. This order matters because teams often retire session storage before they know which fields must survive between calls.
Common mistake: Treating the workflow ID as a permanent identity for the conversation. The ID should resume a bounded workflow, not become a general-purpose token for arbitrary follow-on actions.
Practitioner takeaway: The goal is not to preserve old session behavior in a new wrapper, but to make continuation explicit, bounded, and independently verifiable on every request.