In March 2026, security start-up CodeWall disclosed that its autonomous offensive AI agent had broken into Lilli, McKinsey's internal generative AI platform, in about two hours and without any credentials. According to CodeWall, the agent found publicly exposed API documentation, 22 endpoints that required no authentication and a SQL injection flaw in how one of them handled JSON field names, which gave it read and write access to the production database. CodeWall says that database held 46.5 million chat messages, metadata for 728,000 files, 57,000 user accounts and 95 system prompts. This was a research disclosure, not a criminal breach: CodeWall reported the flaw on 1 March, McKinsey fixed it within hours, and McKinsey says a third-party forensics review found no evidence that client data was accessed by the researcher or anyone else. The identity lesson is simple: an AI platform is only as secure as the authentication and authorisation on the APIs behind it, and the prompts that control its behaviour are writable configuration that needs protecting like any credential.
Key takeaways
- CodeWall's agent found the flaw on 28 February 2026, CodeWall disclosed it to McKinsey on 1 March, McKinsey patched on 2 March and CodeWall published on 9 March 2026.
- Entry point: 22 unauthenticated API endpoints out of more than 200 documented publicly, one of which concatenated JSON keys straight into SQL. No password, token or API key was needed.
- The "46 million chats" figure comes from CodeWall, which says 46.5 million chat messages were stored in plaintext and accessible. Commentators noted CodeWall published no proof of how much was actually retrieved, and McKinsey says it found no evidence that client data was accessed.
- CodeWall says the same injection gave write access to 95 system prompts, so an attacker could have changed Lilli's behaviour "Silently. No deployment needed."
- Lesson: most of this chain was ordinary application security. AI platforms need authenticated APIs, object-level authorisation, parameterised queries and tamper-evident control of prompts and retrieval settings.
At a glance
| Organisation | McKinsey & Company (Lilli, its internal generative AI platform) |
|---|---|
| When | Flaw found 28 February 2026; disclosed to McKinsey 1 March 2026; patched 2 March 2026; published 9 March 2026 |
| Researchers | CodeWall, using an autonomous offensive AI agent, under McKinsey's public responsible disclosure policy |
| Entry point | Publicly exposed API documentation and 22 unauthenticated endpoints, one with SQL injection through JSON keys, chained with an IDOR flaw |
| Identities abused | None were needed: the endpoints accepted anonymous requests. The injection ran with the application's own database privileges, which included writing to system prompts, and reached records for 57,000 user accounts and 384,000 AI assistants (CodeWall figures) |
| Impact | CodeWall says 46.5 million chat messages, 728,000 file records, 57,000 accounts and 95 writable system prompts were exposed; McKinsey says it found no evidence client data or client confidential information was accessed |
| Category | LLM / AI platform; AI agent used as the tester; NHI (unauthenticated APIs and application-level database access) |
What happened
Lilli is McKinsey's internal AI platform, launched in 2023. CodeWall says more than 70% of the firm's 43,000 or more employees use it and that it processes over 500,000 prompts a month. The Register and Inc. put usage at 72% of more than 40,000 staff.
CodeWall, a start-up building offensive AI agents, pointed its agent at the open internet. According to CodeWall's write-up, "The CodeWall research agent autonomously suggested McKinsey as a target citing their public responsible disclosure policy (to keep within guardrails) and recent updates to their Lilli platform." The Stack reported that CodeWall founder Paul Price said McKinsey "approved and proof read" the blog post before it was published.
The agent began by mapping the attack surface. CodeWall says it "found the API documentation publicly exposed", covering more than 200 endpoints, and that 22 of them required no authentication. One of those unprotected endpoints wrote user search queries to the database. The query values were parameterised, but the JSON field names were concatenated directly into SQL. CodeWall says standard scanners, including OWASP ZAP, did not flag this. The agent noticed the field names being reflected in database error messages and, over 15 blind iterations, worked out the query structure until it could read production data.
CodeWall says the agent then chained the injection with an insecure direct object reference (IDOR) flaw to read individual employees' search histories. Within two hours, it says, the agent had full read and write access to the production database. The Stack reported the exercise cost about $20 in tokens.
CodeWall lists what that database contained: 46.5 million chat messages stored in plaintext, 728,000 file records (192,000 PDFs, 93,000 spreadsheets, 93,000 presentations and 58,000 Word documents), 57,000 user accounts, 384,000 AI assistants, 94,000 workspaces, 3.68 million RAG document chunks and 95 system prompt configurations across 12 model types. It also reports more than 266,000 OpenAI vector stores, 1.1 million files and 217,000 agent messages linked to external APIs. For files, CodeWall's evidence is metadata: "The filenames alone were sensitive and a direct download URL for anyone who knew where to look." FStech, reporting the Financial Times, described the haul as chat messages, user account data and the names of hundreds of thousands of PDFs and spreadsheets.
The finding CodeWall stresses most is the prompt layer. Lilli's system prompts were stored in the same database, and the injection had write access. "An attacker with write access through the same injection could have rewritten those prompts. Silently. No deployment needed," CodeWall wrote. Inc. quoted the researchers: "No deployment needed. No code change. Just a single UPDATE statement wrapped in a single HTTP call." CodeWall does not say it altered any prompt.
CodeWall emailed McKinsey's security team on 1 March 2026. The next day, it says, McKinsey's CISO acknowledged the report, and McKinsey patched all unauthenticated endpoints, took the development environment offline and blocked the public API documentation. McKinsey told The Stack: "We promptly confirmed the vulnerability and fixed the issue within hours." It added that its investigation, "supported by a leading third-party forensics firm, identified no evidence that client data or client confidential information were accessed by this researcher or any other unauthorized third party."
Timeline
| Date | Event |
|---|---|
| 2023 | McKinsey launches Lilli internally. |
| 28 February 2026 | CodeWall's agent identifies the SQL injection and confirms the full attack chain, documenting 27 findings. |
| 1 March 2026 | CodeWall sends a responsible disclosure email to McKinsey's security team. |
| 2 March 2026 | McKinsey's CISO acknowledges the report; McKinsey patches the unauthenticated endpoints, takes the development environment offline and blocks the public API documentation. |
| 9 March 2026 | CodeWall publishes its write-up; The Register reports it the same day. |
| 10 March 2026 | Promptfoo and independent analysts question how much of the chain was an AI problem and how much of the data was actually retrieved. |
| 12 to 13 March 2026 | The Stack publishes McKinsey's statement on the third-party forensics review; the Financial Times covers the incident, as reported by FStech. |
How it happened: the identity attack path
- Reconnaissance from public documentation. Full API documentation for more than 200 endpoints was reachable from the internet, giving the agent a map of the application.
- No identity required. 22 endpoints accepted requests with no authentication at all. The attacker never needed to steal or forge a user credential, API key or token.
- SQL injection through JSON keys. One anonymous endpoint that stored search queries built SQL using the names of JSON fields. Error messages leaked enough for the agent to refine its payload over 15 attempts.
- Borrowed database privilege. The injected queries ran with the application's own database access, which covered chat history, user records, file metadata, RAG content and system prompts, with write permission.
- Cross-user access through IDOR. Chaining the injection with an insecure direct object reference let the agent read other employees' search histories, according to CodeWall.
- Control over AI behaviour. Because prompts were rows in a writable table, CodeWall says one UPDATE statement could have changed what Lilli told every user, with no code deployment to review.
Impact
- Exposure, as reported by CodeWall: 46.5 million chat messages, 728,000 file records, 57,000 user accounts, 384,000 AI assistants, 94,000 workspaces, 3.68 million RAG chunks and 95 system prompts. These are CodeWall's figures and have not been independently verified.
- Accessed versus exposed: analyst Edward Kiledjian noted that CodeWall published "no proof-of-concept payloads, no hashes, no screenshots", and that it is unclear whether the figures describe data actually retrieved or data that was reachable. CodeWall's file evidence is filenames and download paths rather than file contents.
- McKinsey's position: McKinsey says a third-party forensics firm found no evidence that client data or client confidential information was accessed by the researcher or any other unauthorised party.
- No known malicious use: none of the sources we reviewed reports exploitation by anyone other than CodeWall.
What this means for NHI governance
The McKinsey case is a reminder that the first identity control on an AI platform is whether its APIs check identity at all. The most serious step in the chain was not a jailbreak or prompt injection: it was 22 endpoints that let anyone call them. Promptfoo's analysis put it plainly, describing the chain as "exposed API surface, missing authentication, unsafe SQL construction, and broken object-level authorization", and concluding that "the model became the interface to a compromised application". Every API behind an AI assistant is called by software, and each call should carry a verified identity and be authorised against the specific object it touches.
The second lesson is about the application's own identity. Once injected, the attacker's queries inherited whatever the application's database account could do. That account could read all users' chats and write the system prompts that govern the model. Separating those privileges, so that the search endpoint's identity cannot touch prompt configuration, would have limited the damage even with the injection in place.
Third, prompts, retrieval rules and assistant configurations are now security-relevant configuration. As Promptfoo notes, "database write access can change model behavior without a code deploy." They deserve the same treatment as secrets and infrastructure code: change control, integrity monitoring and an audit trail of who changed what.
Finally, the tester was itself an AI agent that chose its own target and ran without human intervention, according to The Register. Defenders should assume attackers will use the same approach, which makes basic API hygiene and fast detection more urgent, not less.
Recommendations
- Authenticate every API behind an AI platform. Inventory endpoints, remove anonymous access and do not publish internal API documentation to the internet. The guide to securing enterprise AI copilots and assistants covers the surrounding controls.
- Authorise at the object level. Check that the caller may read each specific chat, search history or document, not just that they are logged in. Retrieval should follow the user's own permissions, as set out in the Permission-Aware RAG Guide.
- Split application database identities. Give each service the least privilege it needs, and keep write access to prompts and model configuration on a separate, tightly held identity.
- Treat prompts as controlled configuration. Version them, require review for changes, hash or sign them and alert on any direct database edit.
- Test AI applications like any other application. Include injection through field names, IDOR and unauthenticated routes in testing, and consider agent-driven testing, as described in Red Teaming AI Agents for Identity Abuse.
- Reduce what sits in plaintext. Set retention limits for chat history and encrypt sensitive conversation data so that a single query cannot return years of it.
Frequently asked questions
Was McKinsey's Lilli AI platform hacked?
Security researchers at CodeWall used an autonomous AI agent to break into Lilli in February 2026 and disclosed the flaw to McKinsey on 1 March 2026 under its responsible disclosure policy. It was a research disclosure rather than a criminal attack, and McKinsey says it found no evidence that client data was accessed by the researcher or anyone else.
Were 46 million McKinsey chats exposed?
The figure comes from CodeWall, which says its agent could access 46.5 million chat messages stored in plaintext. McKinsey has not confirmed the number, and analysts noted CodeWall did not publish evidence of how much data was actually retrieved.
How did the AI agent get into Lilli?
According to CodeWall, it found public API documentation, used one of 22 unauthenticated endpoints and exploited SQL injection in the way JSON field names were built into queries, then chained an IDOR flaw to read other users' data.
Related NHI Mgmt Group resources
OmniGPT breach · DeepSeek breach · McDonald's McHire AI chatbot exposure · EchoLeak Microsoft 365 Copilot · Agentic AI Security Guide
How NHI Mgmt Group can help
Enterprise AI platforms depend on APIs, service identities and configuration that are easy to leave unauthenticated or over-privileged. Our NHI Foundation Level Training Course helps teams govern the non-human identities and AI agents behind these systems.
References
- CodeWall: How We Hacked McKinsey's AI Platform (9 March 2026)
- The Register: AI vs AI: Agent hacked McKinsey's chatbot and gained full read-write access in just two hours (9 March 2026)
- The Stack: McKinsey's AI agent "Lilli" hacked, by another AI agent (12 March 2026, updated 17 March 2026)
- Inc.: An AI Agent Broke Into McKinsey's Internal Chatbot and Accessed Millions of Records in Just 2 Hours (10 March 2026)
- FStech: McKinsey working to fix flaws in AI system after hack (13 March 2026)
- Promptfoo: McKinsey's Lilli Looks More Like an API Security Failure Than a Model Jailbreak (10 March 2026)
- Edward Kiledjian: CodeWall says it hacked McKinsey's AI platform. Here's what holds up, and what doesn't (10 March 2026)