By NHI Mgmt Group Editorial TeamBased on Orca Security: “Pickle in the Pipeline: Critical RCE Vulnerabilities in SGLang’s LLM Serving Framework” (March 12, 2026)

TL;DR: Three unsafe deserialization flaws in SGLang, including two unauthenticated remote code execution paths that trigger when exposed multimodal or disaggregation features accept network input, plus a third crash-dump replay issue tied to malicious .pkl files, were found by Orca Security. The broader lesson is that AI serving frameworks still treat untrusted bytes as trusted control flow, which makes runtime trust boundaries the real security control.


At a glance

What this is: This is a security analysis of three SGLang deserialization vulnerabilities, including two unauthenticated network RCE paths and one malicious crash-dump replay issue.

Why it matters: It matters because AI serving frameworks can turn ordinary network exposure or file handling into code execution, which changes how teams govern AI workload isolation, service access, and deployment trust boundaries.


Context

SGLang is an LLM serving framework, and the core problem here is not model quality but trust handling at the boundary between network input and execution. When a serving stack deserializes untrusted bytes with pickle, it is effectively allowing data to become instructions.

For IAM and platform teams, the lesson is that AI serving infrastructure creates its own identity and trust surface. Network reachability, service-to-service communication, and file-based replay workflows all become security boundaries that need explicit control, not implied trust.

This pattern is typical of fast-moving AI infrastructure: functionality is added for throughput and operability before the trust model is tightened. In that environment, deserialization bugs become deployment risks, not just code defects.


Key questions

Q: What breaks when AI serving frameworks deserialize untrusted network data?

A: The trust boundary collapses before the application can validate the request. If the framework uses pickle or another executable object format, the payload can invoke attacker-chosen functions during parsing, which turns a message handler into a code execution path. In practice, this can expose model data, credentials, and adjacent workloads. The right control is to remove executable deserialization from exposed interfaces and use schema-based parsing instead.

Q: Why do exposed AI brokers create remote code execution risk?

A: Because the broker often receives traffic before any application-level policy check and passes it directly into a deserializer. If the socket is reachable and the format can trigger code execution, network access becomes enough to compromise the process. The core issue is unsafe trust in internal control traffic.

Q: What are the signs that an AI service has already been abused through deserialization?

A: Look for unexpected shell processes, unusual file writes under the AI runtime account, and outbound connections that do not match normal inference traffic. Those indicators suggest a payload executed inside the process rather than failing safely. At that point, the service should be treated as compromised, not merely misconfigured.

Q: Should teams handle replay utilities like production attack surfaces?

A: Yes. Any utility that loads serialized objects can execute attacker-controlled code if the file origin is not tightly governed. Replay scripts, debugging helpers, and crash-dump tools should be restricted, audited, and tested with the same caution as live services whenever they consume .pkl or similar formats.


Technical breakdown

Why pickle turns network input into code execution

Python pickle is not a data-only format. It stores instructions for reconstructing objects, including the callable to invoke and the arguments to pass. If a service calls pickle.loads() on attacker-controlled input, it is delegating execution to the payload itself. That is why the unsafe deserialization issue in SGLang is structurally severe: the broker does not need a second-stage bug, because the deserializer is already an execution primitive. In AI serving systems, this is especially dangerous when the message path is exposed to the network and the application assumes the bytes are internal control traffic.

Practical implication: Treat any pickle-based parser on network input as an execution boundary, not a data parser.

How SGLang exposed the broker path

The vulnerable multimodal and disaggregation paths start a background broker, bind it to tcp://* by default, and then pass received payloads straight into pickle.loads(). That means any reachable host can send a message that triggers deserialization before application validation, authentication, or authorization can happen. The architecture fails because the transport endpoint and the deserializer are coupled too tightly. In practical terms, the broker is not just listening for messages. It is listening for code-shaped input on an unauthenticated socket.

Practical implication: Restrict broker reachability to trusted local or segmented networks before enabling these features.

Why crash-dump replay utilities widen the attack surface

The crash-dump replay script shows a different but related failure mode: local file trust. A utility intended to replay .pkl crash dumps loads them with pickle.load() without validating provenance or content. That turns a convenience workflow into a code execution path when a malicious file lands in a shared directory or is supplied by social engineering. This is a common AI infrastructure pattern. Operational scripts inherit the same deserialization risk as live services, but teams often treat them as benign because they sit in scripts or playground directories rather than production code.

Practical implication: Treat replay and debug utilities as production-grade attack surfaces whenever they consume serialized objects.


Threat narrative

Attacker objective: The attacker aims to execute arbitrary code inside the SGLang process and use that foothold to reach model, credential, or cluster resources.

  1. Entry occurs when an attacker reaches the exposed ZMQ broker on the network or gets an operator to load a malicious crash-dump .pkl file.
  2. Credential harvesting is unnecessary because pickle payloads execute during deserialization, giving the attacker code execution at the moment of load.
  3. Escalation follows immediately as the SGLang process runs attacker-supplied commands in the application context, with possible access to model data, API credentials, and cluster-adjacent assets.
  4. Impact is arbitrary code execution inside the AI serving stack, which can be used to pivot into surrounding infrastructure or exfiltrate sensitive inference workloads.
  • LiteLLM MCP auth bypass 2026: An exploited LiteLLM MCP auth bypass and default sk-1234 master keys let attackers steal AI gateway master and provider API keys.
  • JADEPUFFER agentic ransomware 2026: The first documented agentic ransomware used harvested keys, default MinIO credentials and a default Nacos signing key to wipe a database.

Read and download The State of NHI & AI Agent Breach Report 2026, covering 200+ breaches impacting Non-Human Identities including AI Agents.


NHI Mgmt Group analysis

Untrusted deserialization is an identity boundary failure, not just a coding bug. When an AI serving stack accepts bytes over the network and turns them into live objects with pickle, it collapses the distinction between transport and execution. That means the trust decision happens too late, after the payload has already crossed into process context. Practitioners should treat deserialization as an authorization boundary for machine traffic, not a utility function.

AI serving frameworks often create a runtime trust boundary that their operators do not explicitly govern. Features such as multimodal brokers, disaggregation channels, and replay scripts are added to support throughput and debugging, but they inherit the same security assumptions as production APIs. In this case, the assumption that internal AI control traffic is safe failed because the network endpoint was reachable and unauthenticated. The implication is that AI workload governance must include transport exposure, not just model access.

Ephemeral execution paths create an identity blast radius that is easy to underestimate. A single pickle-loaded payload can execute inside the SGLang process, then leverage whatever that process can already reach, including model files, secrets, and cluster services. That makes the effective identity of the service account or container more important than the application feature that triggered the load. Teams should evaluate AI serving stacks as privilege-bearing identities, not isolated inference boxes.

Deserialization convenience is a form of security debt that accumulates across AI tooling. The article notes more than 20 pickle-related calls in the codebase, which is a strong signal that the risk is systemic rather than isolated. A codebase can survive one unsafe path for a while, but repeated deserialization shortcuts turn into category-level exposure. The practitioner takeaway is that AI infrastructure needs inventory and governance for serializer choice, just as it does for secrets and network ports.

Runtime trust boundaries are now a first-class control in AI operations. OWASP-NHI and ZT-NIST-207 are relevant here because the serving stack behaves like a machine identity with network reach and local execution power. The secure design question is no longer whether the model is safe, but whether every path that feeds the model process is trusted, authenticated, and constrained. That is where governance effort has to move.

From our research library:

What this signals

Deserialization policy now belongs in AI platform governance. Teams deploying serving frameworks need to decide which serializers are permitted, which ports are reachable, and which operational scripts may consume serialized objects. That governance should be explicit because convenience formats like pickle convert internal trust into external attack surface almost by design.

Runtime trust boundaries should be documented alongside model deployment boundaries. An AI service that exposes broker ports or replay utilities is carrying a machine identity with more power than many teams assume. If the process can read secrets, spawn shells, or reach cluster services, then the deserializer is part of the privilege model and must be controlled that way. AI supply chain inventory becomes relevant because serialization choices belong in the software bill of materials and deployment review.

Untrusted bytes are the new control plane risk in AI stacks. The more AI infrastructure leans on helper scripts, background brokers, and cross-component message passing, the more likely it is that a single parsing mistake becomes a process compromise. Teams should assume that any input path that reconstructs Python objects is a security boundary, and that boundary belongs in design reviews, not incident response.


For practitioners

  • Lock down AI broker exposure Keep ZMQ broker ports off untrusted networks and bind AI-serving control channels to localhost or tightly segmented internal ranges only.
  • Eliminate unsafe deserialization paths Inventory every pickle.loads() and pickle.load() path in AI serving code, scripts, and utilities, then replace them with safer formats such as JSON, msgpack, or Protocol Buffers.
  • Disable unused multimodal and disaggregation features Turn off features that open broker or encoder transfer paths when the deployment only needs text inference, and verify that startup flags match the intended runtime scope.
  • Harden crash-dump replay handling Treat replay utilities as attack surfaces, restrict write access to dump directories, and refuse to load serialized files from shared or externally supplied locations.
  • Monitor AI process behavior for deserialization abuse Watch for unexpected child processes, unusual file creation, and outbound connections from the SGLang runtime as indicators that a payload has already executed.

Key takeaways

  • Unsafe deserialization turned SGLang network input and replay files into execution paths, which is why the issue is a trust-boundary failure rather than a narrow parser bug.
  • The article documents two unauthenticated remote code execution paths and one malicious crash-dump replay issue, showing that the problem is systemic across serving and operational workflows.
  • The practical fix is to remove pickle from untrusted paths, shrink network exposure for broker ports, and govern AI runtime trust boundaries with the same rigor used for privileged infrastructure.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

OWASP Non-Human Identity Top 10, OWASP Agentic AI Top 10 and MITRE ATT&CK address the attack and risk surface, while NIST CSF 2.0 and NIST Zero Trust (SP 800-207) set the governance and control requirements practitioners need to meet.

FrameworkControl / ReferenceRelevance
OWASP Non-Human Identity Top 10NHI-04 — Insecure AuthenticationUnauthenticated broker access lets network input reach code execution paths.
NHI-06 — Insecure Cloud Deployment ConfigurationsThe broker binds to all interfaces by default, creating exposed control-plane reachability.
NHI-02 — Secret LeakageProcess compromise can expose model credentials and other secrets in the serving runtime.
Recommendation — Restrict broker access to authenticated, segmented channels before deserialization can occur. Bind AI service control ports to trusted interfaces and verify deployment exposure in every environment. Assume a deserialization compromise can expose secrets and limit what the AI process can access.
OWASP Agentic AI Top 10ASI02 — Tool MisuseThe serving stack can be driven into unsafe actions through attacker-supplied payloads and helper paths.
Recommendation — Constrain tool-like runtime actions so untrusted inputs cannot trigger unsafe execution paths.
MITRE ATT&CKTA0006;TA0008 — Credential Access; Lateral MovementArbitrary code execution in the AI process can be used to reach secrets and move into adjacent systems.
Recommendation — Map deserialization compromise to credential access and lateral movement controls in detection engineering.
NIST CSF 2.0PR.AA-05 — Access Permissions, Entitlements and AuthorizationsReachable AI control ports and process privileges define whether payloads can become compromise.
Recommendation — Review entitlements for AI runtimes and remove any access path not required for the intended deployment.
NIST Zero Trust (SP 800-207)Enterprise-wide Trust Environment — Enterprise-wide Trust EnvironmentThe article shows why network reachability alone should not imply trust in machine-to-machine traffic.
Recommendation — Treat AI broker traffic as untrusted until it is explicitly authenticated and authorized.

Key terms

  • Unsafe Deserialisation: Unsafe deserialisation happens when an application turns untrusted serialized data back into objects without strict validation. In Python, this can become code execution if the format supports callable reconstruction or hidden execution paths. Authentication code should avoid it entirely when processing cookies, sessions, or request data.
  • Device Trust Boundary: The device trust boundary is the point where an endpoint is considered sufficiently verified to access systems, data, or services. It defines the security line between trusted and untrusted device states, based on posture, identity, integrity, and policy. In practice, it governs whether a device can authenticate, connect, or receive sensitive access.
  • AI Serving Layer: The AI serving layer is the runtime service that accepts prompts, routes requests, and returns model output in production. It matters because this layer often holds network reachability, authentication logic, and access to downstream data or compute, making it a privileged non-human identity boundary rather than a simple application wrapper.
  • Serialization Format: A serialization format is the way software stores or transmits structured objects for later reconstruction. In security analysis, the choice matters because some formats are data-only, while others, such as pickle, can embed execution behavior if they are fed untrusted input.

Deepen your knowledge

NHI governance, agentic AI identity, and machine identity lifecycle are core topics in our NHI Foundation Level course, the industry's only accredited NHI security programme. If you are responsible for identity security strategy or NHI governance in your organisation, it is worth exploring.
NHIMG Editorial Note
Published by the NHIMG editorial team on June 10, 2026.
Updated on October 10, 2026.
NHI Mgmt Group, the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org