TL;DR: Microsoft Copilot could return data from GitHub repositories that had been public only briefly and later made private, because Bing cached those pages and exposed so-called zombie data, according to Lasso Security. The finding shows that repository privacy, secret hygiene, and retrieval permissions are still leaky at the identity layer, not just the app layer.
At a glance
What this is: Lasso Security’s research shows that Copilot could retrieve cached content from GitHub repositories that had already been made private, exposing public-then-private data through Bing’s retained index.
Why it matters: For IAM and NHI teams, the key problem is that access decisions made at publication time can outlive the repository state, so retrieval and caching boundaries need governance alongside authentication and permissions.
By the numbers:
- 100+ Python and Node.js internal packages that could be vulnerable to dependency confusion were discovered.
- 300+ private tokens, keys and secrets to GitHub, Hugging Face, GCP and OpenAI were exposed.
Context
The core problem is that a repository can become private after it has already been indexed, cached, and repackaged by external retrieval systems. In that state, the source platform may no longer show the content, but downstream search and AI systems can still surface it.
For identity and access teams, this is not only a repository-privacy issue. It is a data-retention and retrieval-governance issue that sits between source control, search indexing, and AI answer generation, which means the exposure window can extend beyond the original access decision.
Key questions
Q: What breaks when a repository is made private after it was briefly public?
A: What breaks is the assumption that privacy changes erase prior discoverability. Search engines, caches, and AI retrieval layers can retain historical copies, so content may remain reachable even after GitHub shows a 404. Teams need to treat public-to-private transitions as potential exposure events until cached copies and secrets are reviewed, and they should use the 52 NHI breaches Report to understand how often hidden exposure paths matter.
Q: Why do cached GitHub pages create an identity and access risk?
A: Because cached pages can preserve tokens, keys, package names, and internal code after the source repository becomes private. That turns a content-retention issue into an access-risk issue, since stolen or replayed secrets can be used for authentication, package abuse, or broader lateral movement. The risk is not theoretical when the cached data still answers queries.
Q: How should security teams handle briefly public repositories after they are reclassified?
A: Treat them as exposure events until you have confirmed what was indexed, cached, or copied downstream. Then rotate any secrets, review package namespaces, and verify whether retrieval systems can still surface the data. A privacy change is not the same as full containment.
Q: What is the difference between source repository privacy and retrieval privacy?
A: Source repository privacy controls who can open the live object. Retrieval privacy controls whether search indexes, caches, and AI systems can still return the content after the source changes state. Both matter, because an object can be private at the source and still be functionally visible through retained snapshots.
Technical breakdown
How cached public content becomes retrievable after privacy changes
When a repository is public, search engines can index its pages and store cached snapshots. If the repository later becomes private, the live URL may return 404, but the cached copy can still preserve the earlier content. Retrieval-augmented systems that depend on search indexes can then surface that stale material even when direct access to the source no longer exists. This is why the question is not only who can open the repository now, but what was captured during the brief period it was public. Practical implication: treat temporary public exposure as durable exposure until downstream caches are verified and removed.
Practical implication: verify that search indexes and cached copies are purged when sensitive repositories are reclassified.
Why retrieval-augmented generation can ignore the current repository state
RAG systems do not necessarily query the live source of truth every time. They often answer from indexed or cached content that was captured earlier, then combine it with language-model generation. That means the model can appear to “know” private data even when the source repository is now inaccessible. The technical failure is a mismatch between the authorization state of the source system and the availability state of the retrieval layer. Practical implication: governance must cover the retrieval path, not just the repository permission model.
Practical implication: govern what data can be indexed, retained, and reused by retrieval systems before they answer.
Why secrets in code are especially exposed by zombie repositories
If a repository briefly contains tokens, keys, or package references and later becomes private, the cached snapshot can preserve all of that material. That creates a secondary risk beyond code disclosure: exposed secrets can be reused for access, and internal package names can support dependency confusion. In other words, the cached repository becomes both an information disclosure source and an identity compromise source. Practical implication: classify cached-source exposure as a credential incident, not just a content-leak incident.
Practical implication: rotate or revoke any secrets found in previously public repositories and audit package names for confusion risk.
Threat narrative
Attacker objective: The objective is to recover sensitive code and secrets from repositories that users believe are private, then reuse those credentials or internal details for broader compromise.
- Entry occurred when repositories were public long enough for Bing to index and cache their pages.
- Credential and content access followed when Copilot and cached Bing pages surfaced repository material that was no longer directly reachable on GitHub.
- Impact came from the exposure of private code, tokens, keys, and package names that should have remained inside the organisation boundary.
Breaches seen in the wild
- CoPhish OAuth phishing via Copilot Studio: Datadog showed Copilot Studio agents on a Microsoft domain can front OAuth consent phishing and forward stolen tokens; no victims reported.
- Microsoft Azure OpenAI abuse by Storm-2139: Storm-2139 used API keys leaked by Microsoft customers to hijack Azure OpenAI, bypass safety guardrails and resell access to generate harmful content.
Read and download The State of NHI & AI Agent Breach Report 2026, covering 200+ breaches impacting Non-Human Identities including AI Agents.
NHI Mgmt Group analysis
Zombie repository exposure is an identity governance failure, not just a search problem: Once a repository is indexed while public, its data can outlive the access decision that created it. That means the governance boundary is no longer the repository alone, but the entire retrieval chain that can reconstitute formerly public content. Practitioners should treat source-control privacy as a lifecycle state, not a final control.
Retrieval permissions must be governed separately from source permissions: A repository can be private while a search index, cache, or AI retrieval layer still has usable content. That breaks the assumption that current source permissions describe current data reachability. The named concept here is the retrieval residual risk: content remains operationally accessible after the original object has been reclassified. Security teams need to account for this residual layer in access reviews and incident triage.
Cached secrets turn content leakage into identity compromise: The article’s findings are not limited to code disclosure. Tokens, keys, and package references preserved in a cache create a direct path to credential abuse and dependency confusion. The implication is that repository leakage, NHI exposure, and software supply-chain risk now overlap in the same event class, so IAM and AppSec teams have to share ownership.
Least privilege does not end at the repository boundary: If an AI system can answer from cached or indexed material, the effective audience for that material is wider than the direct GitHub ACL suggests. That makes “private” an incomplete security statement unless indexing, retention, and retrieval are also constrained. Practitioners should evaluate whether their controls cover where data can be replayed, not just where it is stored.
Public-then-private content should be handled as compromised until proven otherwise: The article shows that a short public window can create long-lived exposure through downstream systems. That is a lifecycle problem for NHI governance and for data governance alike. The practitioner conclusion is simple: assume any briefly public repository has already escaped into another control plane.
What this signals
Retrieval residual risk: a repository can be private and still remain reachable through caches, indexes, or AI answer layers. That changes the governance target from source permissions alone to the entire retrieval path, including what data is retained after the original access decision.
The practical lesson for identity teams is that temporary public exposure should be handled as a durable control failure until caches are purged and any embedded secrets are rotated. If your programme stops at repository ACLs, you are governing the wrong layer.
For practitioners
- Audit public-then-private repository history Identify repositories that were public at any point, then made private, and treat them as exposure candidates until downstream cache and index checks are complete.
- Revoke and rotate exposed secrets Rotate any tokens, keys, certificates, or service credentials found in cached repository content, even if the live repository is now private.
- Constrain search and retrieval indexing Review which repositories, paths, and file types are eligible for external indexing so that temporary exposure does not become durable exposure.
- Scan for dependency confusion signals Inventory internal package names and publishing patterns that were visible in cached repositories and verify they are not usable by external package ecosystems.
- Add retrieval-layer incident handling Classify cached repository exposure as a security incident with credential-impact triage, not as a simple content cleanup task.
Key takeaways
- A repository that was public even briefly can leave behind cached content that still exposes private material after the source is reclassified.
- Lasso Security’s research found large-scale exposure across thousands of repositories, hundreds of organisations, and hundreds of secrets.
- Teams need controls for indexing, caching, and retrieval, because privacy at the source does not guarantee privacy in downstream systems.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Non-Human Identity Top 10 and MITRE ATT&CK address the attack and risk surface, while NIST CSF 2.0 sets the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP Non-Human Identity Top 10 | NHI-02 — Secret Leakage | Cached repository snapshots exposed secrets, keys, and tokens after the source became private. |
| NHI-03 — Vulnerable Third-Party NHI | Copilot and Bing acted as external retrieval layers that re-exposed content outside the source repository boundary. | |
| NHI-07 — Long-Lived Secrets | The article shows how secrets can remain useful long after a brief public exposure window closes. | |
| Recommendation — Scan for exposed secrets in cached repositories and revoke any credentials found in retained snapshots. Inventory third-party retrieval paths that can surface NHI data beyond the source system's current permissions. Shorten secret lifetimes and rotate any credential that may have been captured during public exposure. | ||
| NIST CSF 2.0 | PR.AA-05 — Access Permissions, Entitlements and Authorizations | Current entitlements did not prevent historical cached content from being retrieved later. |
| PR.DS-01 — Data-at-Rest is Protected | Cached copies preserved data outside the live GitHub repository, creating a separate protection problem. | |
| Recommendation — Extend authorization reviews to indexing and retrieval paths, not only to live repository access. Apply data protection controls to cached and indexed copies of sensitive repository content. | ||
| MITRE ATT&CK | TA0006;TA0010 — Credential Access; Exfiltration | The exposure path delivered credentials and sensitive content that could support downstream compromise or exfiltration. |
| Recommendation — Map cached-repository exposures to credential access and exfiltration detection so responders triage them as compromise indicators. | ||
Key terms
- Zombie Data: Data that users believe is private, deleted, or no longer reachable, but that still persists in caches, indexes, or downstream retrieval systems. In identity terms, the exposure outlives the source system’s access change, so governance must account for residual machine-visible copies.
- Retrieval Residual Risk: The security risk that content remains reachable after the source object has been reclassified, removed, or made private. It captures the gap between live-system access controls and the behaviour of caches, search indexes, and AI systems that retain earlier copies or embeddings.
- Cached Snapshot: A preserved copy of a page or document stored by a search engine or retrieval system for faster reuse. Cached snapshots can outlive source access changes, which makes them a parallel exposure surface for secrets, code, and other sensitive material.
- Temporary Public Exposure: A brief period in which a repository or file was publicly accessible before being locked down again. Even short-lived exposure can be enough for indexing, copying, or caching, which means the security impact often continues after the source is secured.
Deepen your knowledge
NHI governance, agentic AI identity, and machine identity lifecycle are core topics in our NHI Foundation Level course, the industry's only accredited NHI security programme. If you are building or maturing an IAM programme, it is worth exploring.
Published by the NHIMG editorial team on June 9, 2026.
Updated on October 10, 2026.
NHI Mgmt Group, the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org