Join our Newsletter — 33% off our NHI Course
Home› NHI Breaches› Internet Archive Breach 2024: How an Exposed GitLab…
Breach analysis Incident: 28 Sep 2024

Internet Archive Breach 2024: How an Exposed GitLab Token and Unrotated Zendesk Keys Led to Two Breaches in a Month

← All NHI breaches
By Lalit Choda, NHI Mgmt Group Updated 29 September 2026 7 min read
On this page

In October 2024, the Internet Archive was breached twice through the same set of non-human credentials. The attacker told BleepingComputer that the first breach began with a GitLab configuration file exposed on one of the Archive's development servers, containing an authentication token that let them download source code. That code held further credentials, including for the Archive's database management system, which the attacker used to take the user database. Have I Been Pwned confirmed 31 million records, including email addresses, screen names and bcrypt password hashes. BleepingComputer confirmed that the token had been exposed since at least December 2022. The stolen secrets also included API tokens for the Archive's Zendesk support system. On 20 October, people who had once contacted the Archive began receiving replies sent through its genuine Zendesk account, warning that the organisation had not rotated the exposed keys and that the attacker could read more than 800,000 support tickets dating back to 2018.

Key takeaways

  • An exposed GitLab configuration file with an authentication token let the attacker download Internet Archive source code, according to the attacker.
  • BleepingComputer confirmed the token had been exposed since at least December 2022.
  • Credentials in the source code reached the user database; 31 million records were confirmed by Have I Been Pwned.
  • Unrotated Zendesk API tokens from the same leak were used on 20 October to send emails and reach 800,000+ support tickets.
  • The identity lesson: after a breach, every secret in the stolen code must be rotated, or the attacker keeps the keys.

At a glance

OrganisationInternet Archive (non-profit digital library and Wayback Machine)
WhenUser database taken around 28 September 2024; breach disclosed 9 October; Zendesk abuse 20 October 2024
AttackerUnattributed; a separate DDoS campaign was claimed by SN_BlackMeta
Entry pointA GitLab authentication token in an exposed configuration file on a development server, according to the attacker
Identities abusedA GitLab token; database credentials in source code; Zendesk API tokens
Impact31 million user records stolen; access to 800,000+ support tickets; emails sent from the Archive's Zendesk
CategoryNHI. Incident class: confirmed NHI breach (exposed GitLab token and unrotated API tokens)

What happened

BleepingComputer reported that the attacker contacted it and explained that "the initial breach of Internet Archive started with them finding an exposed GitLab configuration file on one of the organization's development servers, services-hls.dev.archive.org." BleepingComputer confirmed that "this token has been exposed since at least December 2022, with it rotating multiple times since then." The attacker said the source code "contained additional credentials and authentication tokens, including the credentials to Internet Archive's database management system," allowing them to "download the organization's user database, further source code, and modify the site." Security Affairs reported that Have I Been Pwned confirmed the stolen archive held 31 million records, and that Troy Hunt dated the most recent records to 28 September 2024.

The second breach used secrets from the first. On 20 October, BleepingComputer received messages from people who had been sent replies to old removal requests. The attacker wrote: "It's dispiriting to see that even after being made aware of the breach weeks ago, IA has still not done the due diligence of rotating many of the API keys that were exposed in their gitlab secrets." They added: "As demonstrated by this message, this includes a Zendesk token with perms to access 800K+ support tickets sent to [email protected] since 2018." The emails passed DKIM, DMARC and SPF checks, showing they came from an authorised Zendesk server. Some people had uploaded identity documents with removal requests.

BleepingComputer said it had tried repeatedly to warn the Archive that its source code had been stolen through the exposed GitLab token but received no response. Infosecurity Magazine quoted vx-underground: "It appears that the person(s) who compromised The Internet Archive still maintain some form of persistent access and are trying to send a message." A DDoS campaign at the same time was claimed by SN_BlackMeta, a different actor.

Timeline

DateEvent
December 2022The GitLab token is exposed on a development server, according to BleepingComputer.
28 September 2024Most recent timestamp in the stolen user database, likely the date of exfiltration.
9 October 2024The breach of 31 million user records is reported alongside DDoS attacks.
20 October 2024Attacker uses unrotated Zendesk tokens to email people who contacted the Archive.
21 October 2024Security Affairs and Infosecurity Magazine report the second breach.

How it happened: the identity attack path

  1. Token exposed. A GitLab configuration file with an authentication token sat on a public development server.
  2. Source code taken. The token let the attacker download the Archive's source code.
  3. More secrets found. The code held database credentials and Zendesk API tokens.
  4. User data stolen. Database credentials gave access to 31 million user records.
  5. Second breach. Unrotated Zendesk tokens were used weeks later to reach support tickets and send emails.

Impact

  • User data: 31 million records including emails, screen names and bcrypt password hashes.
  • Support data: access to more than 800,000 support tickets since 2018, some with identity documents.
  • Trust: emails sent from the Archive's own Zendesk account.

What this means for NHI governance

This breach shows how one exposed token cascades. A source control token opened the code, the code held database and SaaS tokens, and those tokens opened user data and support systems. The second breach is the sharper lesson: the organisation knew its code had been stolen, but the API keys inside it were not all rotated, so the attacker kept working access.

Secrets should not live in code, and after a code leak every secret in it must be treated as compromised and rotated, starting with the most powerful. See our Leaked Credential Response Playbook and Secrets Management Guide.

Recommendations

  • Keep configuration files off public servers. Block access to .git and config paths on development hosts. See our Secrets Management Guide.
  • Rotate every secret in stolen code. Inventory the secrets in the code and revoke them all. See the Leaked Credential Response Playbook.
  • Move secrets out of source code. Use a vault and inject secrets at runtime. See the CI/CD Pipeline Identity Security Guide.
  • Scope SaaS API tokens. A support integration rarely needs access to every ticket since 2018. See the API Key Management Guide.
  • Act on outside warnings. Have a channel for researchers and journalists to report exposures.

Frequently asked questions

How was the Internet Archive breached in 2024?

According to the attacker, an exposed GitLab configuration file on a development server held a token that let them download source code, which contained further credentials used to take the user database.

Why was the Internet Archive breached a second time?

Zendesk API tokens in the stolen code had not been rotated, so the attacker could still use them weeks later.

How many users were affected?

Have I Been Pwned confirmed 31 million user records. The attacker also had access to more than 800,000 support tickets.

EmeraldWhale 2024 · Misconfigured Git Servers 2026 · Cloudflare Breach 2023 · Leaked Credential Response Playbook · Secrets Management Guide

How NHI Mgmt Group can help

After a breach, the secrets you do not rotate are the ones that come back. We help teams find every credential in stolen code and rotate them in the right order. See our NHI and AI agent security training.

References

Explore further

Sign up to our weekly newsletter — get 33% off our NHI Foundation Level Course

    NHIMG Editorial Note
    Written and reviewed by Lalit Choda, NHI Mgmt Group. Last updated 29 September 2026.
    Based on the public sources listed under References. Details may change as investigations continue.
    NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org