TL;DR: Microsoft Copilot could return data from GitHub repositories that had been public only briefly and later made private, because Bing cached those pages and exposed so-called zombie data, according to Lasso Security. The finding shows that repository privacy, secret hygiene, and retrieval permissions are still leaky at the identity layer, not just the app layer.
Editorial analysis by NHI Mgmt Group, based on content published by Lasso Security: “Wayback Copilot: Using Microsoft's Copilot to Expose Thousands of Private GitHub Repositories”.
By the numbers:
- 100+ Python and Node.js internal packages that could be vulnerable to dependency confusion were discovered.
- 300+ private tokens, keys and secrets to GitHub, Hugging Face, GCP and OpenAI were exposed.
Key questions
Q: What breaks when a repository is made private after it was briefly public?
A: What breaks is the assumption that privacy changes erase prior discoverability.
Q: Why do cached GitHub pages create an identity and access risk?
A: Because cached pages can preserve tokens, keys, package names, and internal code after the source repository becomes private.
Q: How should security teams handle briefly public repositories after they are reclassified?
A: Treat them as exposure events until you have confirmed what was indexed, cached, or copied downstream.
Practitioner guidance
- Audit public-then-private repository history Identify repositories that were public at any point, then made private, and treat them as exposure candidates until downstream cache and index checks are complete.
- Revoke and rotate exposed secrets Rotate any tokens, keys, certificates, or service credentials found in cached repository content, even if the live repository is now private.
- Constrain search and retrieval indexing Review which repositories, paths, and file types are eligible for external indexing so that temporary exposure does not become durable exposure.
Bottom line: A repository that was public even briefly can leave behind cached content that still exposes private material after the source is reclassified.
Explore further
View Full Forum → | NHI Foundation Course → | Our Services → | Read the full analysis →
Zombie data is a governance failure, not a search quirk: once content has been public, organisations often assume privacy changes end the exposure story. This case shows that search caches and copilots can preserve the old access state long after the source repository is locked down. The practitioner conclusion is simple: visibility is not the same as revocation, and the data lifecycle must extend beyond the originating platform.
A few things that frame the scale:
- 300+ private tokens, keys & secrets to GitHub, Hugging Face, GCP, OpenAI, etc. were exposed, according to The State of Secrets Sprawl 2025.
- In the same research, 4.6% of all public GitHub repositories contain at least one hardcoded secret, a signal that code exposure and credential exposure are tightly linked.
A question worth separating out:
Q: Who is accountable when an AI assistant surfaces private code from a cached repository?
A: Accountability usually spans the repository owner, the platform owner, and the team governing indexing or retrieval. The source system may be private, but the cached copy may still be live in another layer. That is why privacy incidents involving code and secrets should be handled as cross-platform access governance failures, not isolated GitHub hygiene issues.
👉 Read our full editorial: Wayback Copilot exposes private GitHub repos through cached data
Zombie repository exposure is an identity governance failure, not just a search problem: Once a repository is indexed while public, its data can outlive the access decision that created it. That means the governance boundary is no longer the repository alone, but the entire retrieval chain that can reconstitute formerly public content. Practitioners should treat source-control privacy as a lifecycle state, not a final control.
A question worth separating out:
Q: What is the difference between source repository privacy and retrieval privacy?
A: Source repository privacy controls who can open the live object. Retrieval privacy controls whether search indexes, caches, and AI systems can still return the content after the source changes state. Both matter, because an object can be private at the source and still be functionally visible through retained snapshots.
👉 Read our full editorial: Wayback Copilot exposes private GitHub repos through cached data