TL;DR: ChromaDB CVE-2026-45829 allows unauthenticated remote code execution through the Python FastAPI server’s auth flow, affecting versions 1.0.0 through 1.5.8 and leaving roughly 73% of internet-exposed instances vulnerable, according to Orca Security. The issue shows why AI retrieval backends must be treated as executable infrastructure, not passive data stores.
Editorial analysis by NHI Mgmt Group, based on content published by Orca Security: “Critical Pre-Auth RCE in ChromaDB Threatens AI Infrastructure”.
By the numbers:
- Approximately 14 million monthly PyPI downloads show ChromaDB's broad footprint in enterprise AI deployments.
- Approximately 73% of internet-exposed ChromaDB instances are running vulnerable versions according to Shodan-based scanning data.
- CVE-2026-45829 was disclosed as a max-severity vulnerability with a CVSS score of 10.0.
Key questions
Q: What fails when a vector database can execute code before authentication?
A: The trust boundary fails because the service can run attacker-controlled code before it verifies who sent the request.
Q: Why do AI retrieval backends create host compromise risk instead of just data exposure risk?
A: Because many of them can resolve artefacts, load model code, or invoke runtime dependencies inside the same process that serves search or retrieval.
Q: What are the signs that an AI service is trusting external artefacts too late?
A: Look for request paths that download, import, or execute model references before authentication or policy checks complete.
Practitioner guidance
- Patch to a non-vulnerable ChromaDB release Move internet-facing deployments to ChromaDB 1.5.9 or later, then verify the running version in every environment rather than assuming the patch reached all hosts.
- Remove public reachability from the Python FastAPI server Restrict access to trusted clients only and stop exposing the Python API server to the public internet, especially where RAG or agentic workflows can reach it.
- Treat external model references as untrusted code Scan model artefacts before runtime execution and block trust_remote_code pathways unless the environment is isolated and explicitly approved for code loading.
Bottom line: The defect is a pre-authentication execution path, which means the issue is about trust order as much as it is about patching.
Explore further
View Full Forum → | NHI Foundation Course → | Our Services → | Read the full analysis →
Pre-auth execution in AI backends turns an access-control failure into an identity failure: This vulnerability works because the server evaluates untrusted model configuration before it verifies the caller. That means the auth boundary is placed after the execution boundary, which is a structural failure of trust sequencing rather than a simple missing control. For practitioners, the lesson is that AI retrieval infrastructure cannot be assumed to be inert data plumbing when it can execute remote code as part of request processing.
A few things that frame the scale:
- 85% of organisations lack full visibility into third-party vendors connected via OAuth apps, according to The State of Non-Human Identity Security.
- Only 1.5 out of 10 organisations are highly confident in their ability to secure NHIs, compared to nearly 1 in 4 for securing human identities.
A question worth separating out:
Q: Who is accountable when a pre-authentication RCE affects an AI service?
A: Accountability usually spans application owners, platform teams, and cloud operators because the failure sits across request handling, deployment design, and network exposure. Governance frameworks should assign ownership for runtime execution paths, not just patching, because the key question is who approved an architecture that can execute untrusted input before authentication.
👉 Read our full editorial: ChromaDB pre-auth RCE exposes AI pipelines to full server compromise
Pre-auth execution in AI backends turns an access-control failure into an identity failure: This vulnerability works because the server evaluates untrusted model configuration before it verifies the caller. That means the auth boundary is placed after the execution boundary, which is a structural failure of trust sequencing rather than a simple missing control. For practitioners, the lesson is that AI retrieval infrastructure cannot be assumed to be inert data plumbing when it can execute remote code as part of request processing.
A few things that frame the scale:
- 85% of organisations lack full visibility into third-party vendors connected via OAuth apps, according to The State of Non-Human Identity Security.
- Only 1.5 out of 10 organisations are highly confident in their ability to secure NHIs, compared to nearly 1 in 4 for securing human identities.
A question worth separating out:
Q: Who is accountable when a pre-authentication RCE affects an AI service?
A: Accountability usually spans application owners, platform teams, and cloud operators because the failure sits across request handling, deployment design, and network exposure. Governance frameworks should assign ownership for runtime execution paths, not just patching, because the key question is who approved an architecture that can execute untrusted input before authentication.
👉 Read our full editorial: ChromaDB pre-auth RCE exposes AI pipelines to full server compromise
Pre-auth execution boundary collapse: The real governance failure is not simply missing authentication, but authentication happening after the server has already accepted and acted on untrusted input. That breaks the assumption that policy checks precede side effects. For AI infrastructure teams, the implication is that request handling order is itself a security control surface, not just an implementation detail.
A few things that frame the scale:
- AI-related credential leaks surged 81.5% year-over-year in 2025, with the surrounding AI infrastructure leaking 5x faster than core LLM providers, according to the State of Secrets Sprawl 2026.
A question worth separating out:
Q: How should teams respond when an AI backend mixes data retrieval with code execution?
A: They should treat the backend as executable infrastructure and reduce the exposed blast radius first. That means removing internet reachability where possible, isolating runtime environments, and separating model ingestion from request processing. The core decision is whether the service is allowed to execute external artefacts at all, not just whether it stores embeddings.
👉 Read our full editorial: ChromaDB pre-auth RCE exposes AI pipelines to full server compromise