In November 2023, researchers at Lasso Security searched GitHub and Hugging Face for Hugging Face API tokens that developers had left in public code, and found 1,681 that still worked. Lasso says the tokens gave it access to 723 organisations' accounts, including Meta (Llama 2), EleutherAI (Pythia) and BigScience Workshop (Bloom). 655 of the tokens had write permission, 77 of them for organisations, which would have let an attacker alter models and datasets that, according to The Register, include models with millions of downloads. Lasso also says it could read more than 10,000 private models. The researchers reported their findings, published them on 4 December 2023, and Hugging Face says it invalidated every token they found. There is no evidence that anyone else used the tokens. The case shows how much trust in the AI supply chain rests on simple, long-lived API tokens.
Key takeaways
- Lasso Security found 1,681 valid Hugging Face API tokens exposed in public GitHub and Hugging Face repositories, published on 4 December 2023.
- The researchers validated each token with Hugging Face's whoami API, which told them its owner, organisations and permissions, then mapped the models and datasets each could reach.
- According to Lasso, the tokens reached 723 organisations, 655 had write permission and the team could access over 10,000 private models and write to 14 heavily downloaded datasets. These are the researcher's own figures.
- This was research, not a criminal attack. Hugging Face says all the tokens were invalidated, and Lasso says Meta, Google, Microsoft and VMware revoked theirs the day they were told. No malicious use has been reported.
- The identity lesson: a personal access token with write rights to a popular model is a supply chain credential, and it needs the same protection as a package publishing token.
At a glance
| Organisations | Hugging Face and its users; 723 organisations whose accounts the tokens reached, including Meta, EleutherAI and BigScience Workshop, according to Lasso Security |
|---|---|
| When | Research carried out in November 2023; published 4 December 2023 |
| Attacker | None known. Found and tested by Lasso Security researchers led by Bar Lanyado, who reported to the affected parties |
| Entry point | Hugging Face API tokens hard-coded in public GitHub repositories and public Hugging Face repositories |
| Identities abused | Hugging Face user access tokens, including tokens with write permission to organisations' model and dataset repositories; deprecated organisation API tokens |
| Impact | Researchers gained read and write access to model and dataset repositories of 723 organisations and read over 10,000 private models; tokens invalidated; no misuse by others reported |
| Category | NHI, LLM / AI platform. Incident class: confirmed NHI breach (live exposed tokens used by researchers to reach private models; no malicious use reported) |
What happened
Hugging Face is the main public hub for AI models and datasets, with more than 500,000 models and 250,000 datasets at the time, according to Lasso. Developers use its API tokens to download private models, push new versions and manage organisation repositories. Like any API key, a token pasted into a script and pushed to a public repository can be found and used by anyone.
Lasso's team, led by researcher Bar Lanyado, searched GitHub code with a regular expression for the token format and searched Hugging Face for the token prefix. Both searches returned only a limited number of results per query, so they lengthened the prefix to narrow each search and worked through the results. They checked every candidate with Hugging Face's whoami API, which reported whether a token was valid, who owned it, which organisations it belonged to and what permissions it had. The result was 1,681 valid tokens. Lasso says these gave access to 723 organisations, that 655 tokens had write permission (77 of them for organisations) and that the team gained full read and write access to repositories belonging to Meta Llama 2, BigScience Workshop and EleutherAI. "The gravity of the situation cannot be overstated," Lanyado wrote.
Lasso described three risks. With write access an attacker could replace a popular model with a poisoned one, a supply chain attack on everyone who downloads it. They could tamper with training datasets; Lasso says it could have modified 14 datasets with tens or hundreds of thousands of downloads a month. And they could take private models, more than 10,000 of which Lasso says it could access. The team also found that organisation API tokens, which Hugging Face had deprecated, could still read private models through a small change to the login function in Hugging Face's library, though writing was blocked. Hugging Face fixed that.
Hugging Face's chief executive, Clement Delangue, told The Register: "The tokens were exposed due to users posting their tokens in platforms such as the Hugging Face Hub, GitHub, and others." He added: "All Hugging Face tokens detected by the security researcher have been invalidated". In a statement to TechTarget, Hugging Face said: "In general, we recommend users do not publish any tokens to any code hosting platform." EleutherAI's executive director, Stella Biderman, told The Register: "We are always grateful to ethical hackers for their important work identifying vulnerabilities in the ecosystem". Lasso did not give dates for its reports to Hugging Face and the affected organisations.
Timeline
| Date | Event |
|---|---|
| November 2023 | Lasso Security searches GitHub and Hugging Face, validates 1,681 exposed tokens and maps what each can reach. |
| 2023 | Before publication, Lasso notifies affected users, organisations and Hugging Face; Lasso says Meta, Google, Microsoft and VMware revoke their tokens the same day. |
| 4 December 2023 | Lasso publishes its research; The Register reports it with comment from Hugging Face, which says all detected tokens have been invalidated. |
| 5 December 2023 | TechTarget reports the findings with Hugging Face's advice not to publish tokens to code hosting platforms. |
How it happened: the identity attack path
- Tokens committed to public code. Developers hard-coded Hugging Face API tokens in scripts, notebooks and configuration files and pushed them to public GitHub and Hugging Face repositories.
- Search by token format. Because the tokens share a recognisable prefix, simple code searches found them in bulk, with no need to target any individual organisation.
- Validation and reconnaissance through the API. The whoami API told the researchers which tokens were live and exactly which organisations and permissions each carried, turning a list of strings into a map of access.
- Use of the access. Valid tokens gave read access to private models and, for 655 tokens, write access to model and dataset repositories, including those of major AI developers.
- Legacy tokens still working. Deprecated organisation tokens could still read private models through a modified login call, showing that retiring a credential type is not the same as disabling it.
Impact
- Confirmed (research access): Lasso says it reached 723 organisations' accounts, held 655 tokens with write permission and could access more than 10,000 private models. All detected tokens were later invalidated.
- Not confirmed: no model or dataset was altered and no malicious use of the tokens by anyone else has been reported.
- Potential: poisoned models or datasets distributed to everyone who downloads them, theft of private models, and abuse of organisations' accounts, according to Lasso's analysis.
What this means for NHI and AI agent security
AI models are software that organisations and agents pull in automatically, and the right to publish a new version is held by an API token. This case belongs on an NHI list because those tokens were the whole story: no vulnerability was needed, only tokens left in public code and an API that confirmed what each could do. Several were held by individual developers but carried write access to organisation repositories whose models have millions of downloads, so one person's careless commit could have become a supply chain attack on the whole ecosystem.
The fix is the same as for package registries. Tokens with publishing rights should be scoped to one repository, short-lived and kept out of code, and platforms should scan for leaked tokens and revoke them automatically. Hugging Face later suffered its own Spaces secrets breach in 2024, and the same pattern of secrets left in published artefacts appears in our PyPI secrets exposure page. Our AI Supply Chain and AI-BOM Guide and API Key Management Guide cover the controls.
Recommendations
- Revoke and replace any token that has appeared in code. Assume it has been found. Rotate it, remove it from the repository and its history, and review what it was used for. See the Leaked Credential Response Playbook.
- Scope tokens to the least they need. Use read-only tokens wherever possible and fine-grained tokens limited to specific repositories for publishing, never a broad write token in a notebook. See our API Key Management Guide.
- Keep tokens out of code and notebooks. Load them from environment variables or a secrets manager, and use pre-commit hooks and repository scanning to block them. See our Secrets Management Guide.
- Govern who can publish your models. Review which people and tokens hold write access to organisation repositories on Hugging Face and remove access for those who do not need it. See our AI Supply Chain and AI-BOM Guide.
- Verify models before you use them. Pin model versions by commit, check integrity and prefer safe serialisation formats, so a poisoned upload does not flow straight into production or agents.
- Disable deprecated credential types, do not just deprecate them. Lasso's finding on organisation tokens shows legacy credentials can keep working long after they are retired.
Frequently asked questions
What did Lasso Security find on Hugging Face?
Lasso Security found 1,681 valid Hugging Face API tokens exposed in public GitHub and Hugging Face repositories. According to Lasso, they gave access to 723 organisations including Meta, EleutherAI and BigScience Workshop, 655 had write permission, and the team could read more than 10,000 private models.
Were Meta's Llama models compromised through exposed Hugging Face tokens?
Lasso says it obtained full read and write access to Meta Llama 2 repositories through exposed tokens, but there is no report that any model was changed. The researchers reported the tokens, Lasso says Meta revoked its tokens the same day, and Hugging Face invalidated all the tokens found.
How can I protect Hugging Face access tokens?
Never put tokens in code, notebooks or public repositories. Use read-only or fine-grained tokens scoped to specific repositories, store them in environment variables or a secrets manager, scan repositories for leaks, and revoke any token that has been exposed straight away.
Related NHI Mgmt Group resources
Hugging Face Spaces breach 2024 · PyPI secrets exposure 2023 · xAI API key leak 2025 · AI Supply Chain and AI-BOM Guide · API Key Management Guide
How NHI Mgmt Group can help
AI platform tokens are spreading through notebooks, scripts and pipelines faster than most security teams can track them. We help organisations find the tokens that can publish or read their models, scope them properly and build detection and revocation into daily practice. See our NHI and AI agent security training.
References
- Lasso Security (Bar Lanyado): +1500 HuggingFace API Tokens were exposed, leaving millions of Meta-Llama, Bloom, and Pythia users vulnerable (4 December 2023)
- The Register: Exposed Hugging Face API tokens offered full access to Meta's Llama 2 (4 December 2023)
- TechTarget: Exposed Hugging Face API tokens jeopardized GenAI models (5 December 2023)