TL;DR: The Unlimited Technology Systems breach exposed 3,803,750 people and showed how scanned IDs, insurance cards, and intake forms remain the hardest regulated data to inventory, according to Sentra. The lesson is that visibility, retention, and access governance now define breach impact more than the intrusion path itself.
NHIMG editorial — based on content published by Sentra covering the Unlimited Technology Systems data breach: how scanned healthcare records created a hidden exposure risk
Questions worth separating out
Q: What fails when scanned documents are not included in data discovery?
A: The organisation loses visibility into the most sensitive part of its file estate.
Q: Why do scanned healthcare records create more governance risk than structured fields?
A: Structured fields can be inventoried, queried, and controlled with predictable rules.
Q: How do security teams know whether image-based PHI is actually governed?
A: They should be able to identify every repository that stores scans, prove that OCR or equivalent classification runs on those files, and show which identities can access them.
Practitioner guidance
- Inventory scanned-document repositories Locate every file share, fax gateway, portal bucket, and document management system that stores image-based PHI or identity documents, then verify whether discovery tooling can read them end to end.
- Classify image-based files before retention expires Apply OCR and content classification to scanned IDs, insurance cards, and intake forms so retention policies can remove documents collected for one-time verification instead of keeping them indefinitely.
- Scope access to named identities Review which human users, service accounts, and AI agents can reach repositories that contain scanned documents, then remove broad access and document the business justification for each exception.
What's in the full article
Sentra's full analysis covers the operational detail this post intentionally leaves for the source:
- Step-by-step OCR and classification workflow for scanned PDFs, images, and other unstructured files
- Practical guidance for identifying hidden PHI in repositories that standard discovery tools often skip
- Repository-level examples showing how continuous classification supports retention and exposure analysis
- Operational detail on how to handle cloud, SaaS, and on-premises unstructured content inside your environment
👉 Read Sentra's analysis of the Unlimited Technology Systems healthcare breach →
Scanned documents in healthcare: what identity teams are missing?
Explore further
Scanned-document blindness is now a governance failure, not just a discovery gap. The breach shows that organisations can hold highly sensitive identity evidence without being able to enumerate it. That is a control failure because regulators and incident responders care about what was exposed, not whether the file was easy to classify. Practitioners should treat image-based repositories as regulated data stores, not miscellaneous content.
A question worth separating out:
Q: Who is accountable when a business associate exposes scanned identity documents?
A: The business associate is responsible for protecting the data it processes, but the covered entity still needs contractual clarity about data types, retention, and access scope. Accountability also extends to internal owners who sent the files downstream without defining how long scans should live and who should be able to retrieve them.
👉 Read our full editorial: Scanned healthcare records expose a blind spot in data security