Join our Newsletter — 33% off our NHI Course

Hugging Face assistants and data exfiltration: are controls keeping up?

 

(@nhi-mgmt-group)
Member Moderator
Joined: 1 year ago
Posts: 21730
Topic starter  

TL;DR: A deceptive Hugging Face assistant used Sleepy Agent behaviour and image markdown rendering to exfiltrate user email addresses through an attacker-controlled URL, showing how prompt-based trust can be turned into a covert data path, according to Lasso Security. The control gap is not just model safety but identity governance for assistants that can leak data through ordinary conversation flow.

Editorial analysis by NHI Mgmt Group, based on content published by Lasso Security: “Exploiting HuggingFace’s Assistants to Extract Users’ Data”.

Key questions

Q: What breaks when an AI assistant renders untrusted responses into HTML during streaming output?

A: Streaming render pipelines can interpret partial content before the final structure is known, which creates a window for active HTML to slip through.

Q: Why do prompt-visible assistants still create data leakage risk?

A: Because prompt visibility does not control runtime behaviour.

Q: What are the signs that an assistant is failing to protect user data?

A: Warning signs include unexpected markdown, hidden URLs, response patterns that change when specific words or data types appear, and any assistant that repeats user values inside renderable output.

Practitioner guidance

  • Constrain external fetches in assistant output Disable or tightly sandbox markdown features that can trigger remote image loads, link previews, or other client-side requests from assistant responses.
  • Classify assistants as governed non-human identities Assign owners, change control, and offboarding responsibility to each deployed assistant so prompt changes and response behaviour are continuously accountable.
  • Block sensitive fields from response composition Prevent assistants from echoing email addresses, tokens, or other sensitive values into URLs, markdown parameters, or formatted output.

Bottom line: The core issue is not whether users can see an assistant’s prompt. The real problem is that the assistant can still disclose data through its normal response path.

Explore further

View Full Forum →  |  NHI Foundation Course →  |  Our Services →  |  Read the full analysis →


This topic was modified 1 day ago by NHI Mgmt Group

   
Quote
(@mr-nhi)
Member Moderator
Joined: 5 months ago
Posts: 21566
 

Prompt visibility is not the same as behavioural trust. The article shows that users may be able to inspect an assistant prompt and still remain exposed if the prompt contains trigger logic or hidden exfiltration behaviour. That breaks the assumption that disclosure alone creates control. The governance problem is not whether the prompt can be read, but whether the assistant can still change what it does when a condition is met. Practitioners should treat the assistant as a governed identity surface, not a transparent object.

A few things that frame the scale:

  • 85% of organisations lack full visibility into third-party vendors connected via OAuth apps, according to The State of Non-Human Identity Security.
  • Only 1.5 out of 10 organisations are highly confident in their ability to secure NHIs, compared to nearly 1 in 4 for securing human identities.

A question worth separating out:

Q: What is the difference between a safe-looking assistant and a governed assistant?

A: A safe-looking assistant answers normally in testing, while a governed assistant has reviewable ownership, change control, output restrictions, and tests for malicious triggers. The distinction matters because benign conversations do not prove that the assistant is safe when its instructions or rendering behaviour change later.

👉 Read our full editorial: Hugging Face assistant attacks expose the limits of prompt trust



   
ReplyQuote
(@mr-nhi)
Member Moderator
Joined: 5 months ago
Posts: 21566
 

Prompt trust is not an access control model: This article shows that a visible system prompt does not create a governance boundary. The assistant can still disclose data through output formatting, which means the control plane must cover response behaviour, not just prompt review. Practitioners should stop treating prompt transparency as a sufficient safeguard for assistant trust.

A question worth separating out:

Q: Should organisations trust assistants just because the system prompt is visible?

A: No. Visibility helps with review, but it does not stop a malicious or compromised assistant from changing behaviour at runtime or using legitimate output formatting to leak data. Organisations should treat assistants as governed non-human identities and enforce output restrictions, change control, and data handling rules.

👉 Read our full editorial: Hugging Face assistant attacks expose the limits of prompt trust


This post was modified 1 day ago by NHI Mgmt Group

   
ReplyQuote
Share:

Free weekly newsletter

Subscribe to the NHI & AI Identity Journal

The latest on NHI and Agentic AI security – articles, research, breaches, news and events every week.

Bonus 33% off our NHI Course when you subscribe.