In February 2025, Truffle Security scanned the December 2024 Common Crawl archive, a free public snapshot of the web used to train many large language models, and found 11,908 live secrets: API keys, passwords and other credentials that still authenticated successfully with their services. The archive covered 400 terabytes of data from 2.67 billion web pages, and the live secrets appeared on 2.76 million of those pages. Most had been hard-coded by developers into front-end HTML and JavaScript. Mailchimp API keys were the most common, with nearly 1,500 unique keys; one WalkScore API key appeared 57,029 times across 1,871 subdomains, and one AWS root key sat in front-end code. Truffle Security worked with the most affected vendors to revoke several thousand keys. Its point was twofold: these credentials are exposed on the public web, and models trained on such code may learn to suggest hard-coding secrets.
Key takeaways
- Truffle Security found 11,908 live secrets in Common Crawl's December 2024 archive, a dataset used to train LLMs including DeepSeek.
- The secrets appeared on 2.76 million web pages, and 63% were repeated across multiple pages.
- TruffleHog detected 219 secret types; Mailchimp API keys were the most common, with nearly 1,500 unique keys hard-coded in front-end code.
- Truffle Security worked with the most affected vendors to rotate or revoke several thousand keys; no misuse of these keys has been reported.
- The identity lesson: a key embedded in a public web page is public, it is copied into archives and training data, and only rotation removes the risk.
At a glance
| Organisations | Thousands of website owners whose pages contained hard-coded secrets; Common Crawl (dataset); Truffle Security (research) |
|---|---|
| When | Common Crawl December 2024 archive scanned; findings published 27 February 2025 |
| Attacker | None known. Found by Truffle Security researchers |
| Entry point | API keys, passwords and webhooks hard-coded into public web pages, then captured by the web crawl |
| Identities abused | 11,908 live secrets across 219 types, including Mailchimp and WalkScore API keys, Slack webhooks and an AWS root key |
| Impact | Live credentials exposed on the public web and in a widely used LLM training dataset; several thousand keys revoked; no confirmed misuse |
| Category | NHI, LLM and AI platform. Incident class: exposure (secrets in public web data and LLM training data, no confirmed misuse) |
What happened
Truffle Security had earlier written about LLMs telling developers to hard-code API keys, and wanted to know whether training data was part of the reason. Popular LLMs, including DeepSeek, are trained on Common Crawl, so the team downloaded the December 2024 archive, "400 terabytes of web data from 2.67 billion web pages," and scanned it with its open-source tool TruffleHog on 20 servers. It counted only secrets that passed TruffleHog's automated verification: "a secret was considered 'live' only if TruffleHog's automated verification process (which includes service-specific authentication checks) confirmed its validity."
The result, published on 27 February 2025, was 11,908 live secrets on 2.76 million web pages, and the researchers noted "High Reuse Rate among secrets: 63% were repeated across multiple web pages." TruffleHog found 219 different secret types. "Nearly 1,500 unique Mailchimp API keys were hardcoded in front-end HTML and JavaScript," the researchers wrote. One page held 17 live Slack webhooks, and one AWS root key had been placed in front-end HTML for S3 basic authentication. Some development firms reused the same key across client sites, making it possible to identify their customers.
Truffle Security stressed that the leak was not Common Crawl's fault: "Common Crawl should not be tasked with redacting secrets." Rather than contact around 12,000 website owners, "We contacted the vendors whose users were most impacted and worked with them to revoke their users' keys," it said, and "We successfully helped those organizations collectively rotate/revoke several thousand keys." BleepingComputer noted that LLM training data goes through pre-processing to filter sensitive content, but that this offers no guarantee of removing it all.
Timeline
| Date | Event |
|---|---|
| December 2024 | Common Crawl captures the web archive that Truffle Security later scans. |
| 27 February 2025 | Truffle Security publishes its findings. |
| 28 February 2025 | Bitdefender reports the research. |
| 2 March 2025 | BleepingComputer reports the research. |
How it happened: the identity attack path
- Secrets hard-coded in web pages. Developers placed API keys and webhooks in front-end HTML and JavaScript instead of keeping them server-side.
- Public exposure. Anyone viewing the page source could read the secrets.
- Captured by the crawl. Common Crawl archived the pages, secrets included.
- Carried into training data. The archive is used to train LLMs, which may learn the insecure pattern.
- Found and partly revoked. Truffle Security verified the secrets and helped vendors revoke several thousand keys.
Impact
- Exposed: 11,908 live secrets across 219 types, on 2.76 million web pages.
- Potential: attackers could use leaked Mailchimp keys for phishing, brand impersonation and data exfiltration, the researchers said.
- Misuse: none reported.
- Indirect: training data full of hard-coded credentials may reinforce insecure coding suggestions from AI assistants.
What this means for NHI governance
A secret in a web page is not a secret. Front-end code is sent to every visitor, copied by crawlers, stored in archives and, as this research shows, used to train AI models. Once that has happened, removing the key from the page does not undo the exposure. The only fix is to revoke and replace it. The high reuse rate also matters: one key used across thousands of pages or client sites means one leak exposes all of them.
There is also a feedback loop for AI-assisted development. If models learn from code that hard-codes keys, they may suggest the same pattern, and developers who accept those suggestions create new leaks. Secret scanning in pipelines, and clear guidance for coding assistants, both help break that loop. See our Secrets Management Guide, API Key Management Guide and AI Coding Agents Security Guide.
Recommendations
- Keep secrets out of front-end code. Call third-party APIs from the server, or use keys designed for public use with strict restrictions. See our API Key Management Guide.
- Scan public web assets, not just repositories. Truffle Security recommends extending secret scanning to public web pages and archived datasets. See our Secrets Management Guide.
- Revoke exposed keys, do not just remove them. Archived copies persist. See the Leaked Credential Response Playbook.
- Use a unique key per site or client. Reuse multiplies the impact of a single leak.
- Guide AI coding assistants. Use assistant rules and pipeline checks to stop suggested code from hard-coding credentials. See the AI Coding Agents Security Guide.
Frequently asked questions
How many secrets were found in LLM training data?
Truffle Security found 11,908 live API keys, passwords and other secrets in Common Crawl's December 2024 archive, which is used to train LLMs including DeepSeek.
Were these secrets used by attackers?
No misuse has been reported. The secrets were publicly visible on web pages, and Truffle Security worked with affected vendors to revoke several thousand of them.
Why do secrets in training data matter?
They show the keys were exposed on the public web and archived. Truffle Security also suggests that training on code with hard-coded secrets may encourage LLMs to suggest insecure code.
Related NHI Mgmt Group resources
iOS Apps Leaking Secrets 2025 · DeepSeek Database Exposure 2025 · Secrets Management Guide · API Key Management Guide · AI Coding Agents Security Guide
How NHI Mgmt Group can help
Keys leak wherever code goes, including web pages, archives and AI training data. We help teams find exposed secrets, revoke them properly and stop new ones being hard-coded. See our NHI and AI agent security training.
References
- Truffle Security: Research finds 12,000 'Live' API Keys and Passwords in DeepSeek's Training Data (27 February 2025)
- Bitdefender: 400 TB Data Set Used to Train AI Has API Keys and Valid Credentials, Researchers Find (28 February 2025)
- BleepingComputer: Nearly 12,000 API keys and passwords found in AI training dataset (2 March 2025)