By NHI Mgmt Group Editorial TeamBased on Lasso Security: “+1500 HuggingFace API Tokens were exposed, leaving millions of Meta-Llama, Bloom, and Pythia users vulnerable” (March 16, 2026)

TL;DR: Exposed API credentials can extend from code repositories into model, dataset, and supply-chain compromise, according to Lasso Security, which found 1,681 valid Hugging Face and GitHub tokens, including 655 with write permissions, and mapped access across 723 organisation accounts. Hard-coded tokens turn LLM platforms into identity-risk amplifiers, not just development tools.


At a glance

What this is: This research shows that exposed Hugging Face and GitHub tokens can grant broad access across model and dataset repositories, turning LLM supply chains into a credential exposure problem.

Why it matters: It matters because IAM and NHI teams need to treat developer tokens, repository access, and model distribution paths as one governed identity surface, not separate risk silos.

By the numbers:

  • Lasso Security found 1,681 valid tokens through Hugging Face and GitHub.
  • The researchers mapped access across 723 organisation accounts.
  • They identified 655 users’ tokens with write permissions.

Context

Exposed API tokens in LLM ecosystems are a governance problem, not just a developer hygiene issue. When a token can read, write, or publish models and datasets, the credential becomes part of the supply chain and can influence downstream users at scale.

Hugging Face is a distribution and collaboration layer for models and datasets, so token exposure creates a blended identity surface across repositories, private assets, and shared model infrastructure. The key question for practitioners is not whether a token was found, but how far its permissions reach once it exists outside governed storage.


Key questions

Q: What breaks when LLM platform tokens are exposed in public code or repositories?

A: Exposed tokens turn repository content into an authentication source. Once validated, they can reveal private models, datasets, and organisation memberships, and in some cases allow write access that changes what downstream users trust. The failure is not only leakage but the collapse of repository boundaries into a broader supply chain trust problem.

Q: Why do write-capable AI platform tokens create more risk than read-only tokens?

A: Read-only tokens can expose sensitive assets, but write-capable tokens can alter them. That shifts the risk from disclosure to integrity compromise, including malicious model publication and dataset poisoning. In LLM ecosystems, integrity failures are especially dangerous because downstream consumers often rely on repository content without independent verification.

Q: How can security teams detect AI credential abuse before it becomes a campaign?

A: Look for abnormal model usage, sudden changes in call volume, repeated access from unfamiliar contexts, and activity that crosses normal session boundaries. Combine those signals with token inventory and ownership data so the investigation can move from a single alert to a credential-level view of likely abuse.

Q: Should organisations treat Hugging Face and GitHub access as the same governance problem?

A: Yes, when the same token can move from code exposure to model or dataset access. The governance issue is the shared identity path, not the platform name. Teams should align token inventory, scope, and revocation rules across both environments so a leak in one does not become a trust failure in the other.


Technical breakdown

How exposed LLM tokens expand repository trust

Hugging Face API tokens are not passive secrets. In this context they acted as bearer credentials that could authenticate a user or organisation and then authorize repository actions such as reading private assets, creating models, or modifying files. Because the platform sits between developers, models, and datasets, one exposed token can bridge multiple trust zones. That makes the token itself a supply-chain control point, not just an access artifact. The technical risk is amplified when tokens are long-lived, broadly scoped, or reused across platforms, because leakage in one environment becomes reachable access in another.

Practical implication: treat model-hosting tokens as high-value NHI credentials and scope them to the minimum repository action required.

Why write permissions change the attack surface

Read access exposes intellectual property and private models, but write access changes the threat from disclosure to manipulation. A token with write permissions can add or alter repository contents, publish malicious models, poison datasets, or overwrite trusted artifacts that other teams consume. In an LLM supply chain, downstream consumers often trust repository content implicitly, so the credential holder inherits the ability to shape what others download and run. That is why the permission model matters as much as the token itself: a token that can publish or modify content can become a propagation mechanism for compromise.

Practical implication: separate read-only access from publish rights and audit every token that can modify models or datasets.

Why token discovery becomes a scaling problem

The article shows a familiar pattern in modern credential exposure: search at scale, then validation at scale. Public code search, substring probing, and API validation let attackers filter huge token pools down to usable credentials quickly. Once validated, those tokens can be mapped to organisations, users, and accessible artefacts, revealing the real blast radius. This is why exposed tokens are not isolated events. They create a discovery-to-abuse pipeline in which leaked credentials can be operationalized rapidly across many repositories and accounts before owners notice.

Practical implication: monitor for public token exposure continuously and assume exposed credentials will be validated and weaponized quickly.


Threat narrative

Attacker objective: The objective is to obtain authenticated control over high-value LLM repositories and use that access to steal, modify, or poison models and datasets.

  1. Entry begins with API tokens exposed in public GitHub and Hugging Face repositories or search results, giving outsiders a valid credential set to test.
  2. Credential access follows when those tokens are validated through platform APIs such as whoami, revealing user identity, organisation membership, and permission scope.
  3. Escalation occurs when write-capable tokens allow repository modification, private model access, or dataset tampering across trusted LLM assets.
  4. Impact is achieved through model theft, data poisoning, and supply chain compromise that can affect downstream consumers of the exposed models and datasets.

Read and download The State of NHI & AI Agent Breach Report 2026, covering 200+ breaches impacting Non-Human Identities including AI Agents.


NHI Mgmt Group analysis

Exposed model-hosting tokens are now a supply chain governance problem, not a narrow secrets issue. When a single bearer token can read private assets, modify repositories, and influence what downstream teams download, it crosses from credential hygiene into supply chain integrity. The governance unit is no longer the repository alone but the full model distribution path. Practitioners should treat model platform credentials as part of the software and AI supply chain.

LLM ecosystems collapse the old boundary between development access and production trust. The article shows that tokens found in code and repositories can reach private models, datasets, and organisation accounts with very different blast radii. That means access governance for AI platforms has to assume that developer convenience can become distribution-scale risk. Identity teams should stop separating repo access from model trust as if they were different problems.

Standing token privilege creates a model integrity exposure window. These credentials persist outside the person, pipeline, or workflow that created them, so their effective lifetime can outlast the intent behind the access grant. The result is a reusable trust artifact that can be rediscovered and replayed long after the original deployment context has changed. The implication is simple: the platform’s security posture is only as strong as its least governed token.

Token exposure turns third-party AI platforms into identity control multipliers. The article’s findings show that one leaked credential can surface multiple organisation memberships and broad repository rights across a shared ecosystem. That is exactly the kind of cross-boundary access pattern that breaks conventional assumptions about isolated application controls. Practitioners should evaluate AI platform credentials with the same discipline they apply to privileged service accounts and external supplier access.

Model theft and dataset poisoning are governance outcomes of weak credential lifecycle control. Once a token can authenticate and then write, the distinction between access and compromise narrows sharply. The important question is not whether the platform is popular, but whether the token lifecycle is bounded, monitored, and revoked before it becomes an attack path. Teams should review AI platform access as a lifecycle issue, not a one-time provisioning exercise.

From our research library:

What this signals

Model distribution now behaves like an identity plane. When a token can create repositories, read private artefacts, and affect downstream users, governance has to track issuance, scope, and revocation as one control surface. Teams that still separate developer secrets from AI platform trust will miss the path by which exposed credentials become supply chain compromise.

Hugging Face token exposure shows why identity reviews must include AI content pipelines. Repository access and model trust are now linked, so the review unit should be the credential, not just the human owner or the hosting platform. Security teams should map which tokens can alter models or datasets and decide whether those rights belong in the same approval path as production access.


For practitioners

  • Restrict model-platform tokens to the minimum required scope Issue separate credentials for read, write, and administrative actions, and avoid reusing the same token across model hosting, dataset access, and automation workflows.
  • Inventory exposed AI platform credentials Scan repositories, issue trackers, and CI logs for Hugging Face and GitHub tokens, then revoke any credential that can reach private models or datasets.
  • Add validation checks before release Make token detection part of code review and pre-merge controls so hard-coded credentials are blocked before they reach public repositories.
  • Separate repository access from distribution trust Review which tokens can create, modify, or publish models and datasets, and treat those privileges as supply chain rights rather than routine developer access.

Key takeaways

  • Exposed Hugging Face and GitHub tokens can move beyond simple secret leakage and become a supply chain integrity issue for LLM platforms.
  • The article shows that valid credentials can reach private models, datasets, and organisation accounts, which expands the blast radius far beyond code repositories.
  • Security teams should govern AI platform tokens as privileged identity assets, with tight scope, continuous discovery, and immediate revocation when exposure is detected.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

OWASP Non-Human Identity Top 10, OWASP Agentic AI Top 10 and MITRE ATT&CK define the specific risk controls and attack patterns relevant to this term.

FrameworkControl / ReferenceRelevance
OWASP Non-Human Identity Top 10NHI-02 — Secret LeakageThe article centers on exposed API tokens leaking through public repositories and search.
NHI-05 — Overprivileged NHISome exposed tokens had write permissions and broad repository reach.
NHI-07 — Long-Lived SecretsLeaked tokens remain useful until explicitly revoked, extending exposure windows.
Recommendation — Scan code and repositories for leaked NHI secrets and revoke exposed tokens immediately. Reduce token scope so NHI credentials cannot modify models or datasets by default. Shorten token lifetime and enforce rapid revocation for any exposed credential.
OWASP Agentic AI Top 10ASI03 — Identity & Privilege AbuseLLM and AI platform tokens can be abused to alter model content and trust paths.
Recommendation — Limit identity scope so AI-related credentials cannot be repurposed for unauthorized modification.
MITRE ATT&CKTA0006; TA0008 — Credential Access; Lateral MovementLeaked tokens enabled credential use and movement from public exposure into private repository access.
Recommendation — Map exposed-token activity to credential access and lateral movement detections.

Key terms

  • LLM Supply Chain: The collection of models, datasets, repositories, tokens, and build paths that determine what an AI system consumes and distributes. In practice, it is a trust chain for AI content, and a compromise at any point can affect downstream users and applications.
  • Bearer Token: A bearer token is a credential that grants access to whoever possesses it, without requiring strong proof that the holder is the intended client. In NHI environments, that makes theft and replay the main risk, especially when tokens are long-lived, broadly scoped, or stored in local files.
  • Model Integrity: Model integrity is the degree to which an AI system’s learned behaviour remains faithful to intended design, training assumptions, and governance boundaries. It extends beyond access control to include data lineage, prompt handling, testing evidence, and ongoing monitoring.
  • Token scope: Token scope is the set of actions and systems a credential can reach. In non-human identity governance, narrow scope limits blast radius, while broad scope lets a stolen token publish packages, access cloud services, or mutate repositories far beyond its intended purpose.

Deepen your knowledge

NHI governance, agentic AI identity, and machine identity lifecycle are core topics in our NHI Foundation Level course, the industry's only accredited NHI security programme. If you are building or maturing an IAM programme, it is worth exploring.
NHIMG Editorial Note
Published by the NHIMG editorial team on June 9, 2026.
Updated on October 10, 2026.
NHI Mgmt Group, the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org