TL;DR: Enterprises cannot govern AI safely without first governing unstructured data, because SharePoint, S3, and Drive content is being pulled into RAG pipelines, fine-tuning sets, and knowledge bases before teams understand what it contains, according to BigID. The governance problem is shifting from output filtering to pre-ingestion control, and that makes data discovery and minimization the decisive control plane.
NHIMG editorial — based on content published by BigID: Governing unstructured data is now the AI risk control point
By the numbers:
- When AWS credentials are exposed publicly, attackers attempt access within an average of 17 minutes and as quickly as 9 minutes in some cases.
- 72% of organisations have experienced or suspect they have experienced a breach of non-human identities, with 46% confirmed and 26% suspected.
Questions worth separating out
Q: How should security teams govern unstructured data for GenAI use cases?
A: Security teams should govern unstructured data by mapping content to business context, human relevance, and downstream AI use paths, not by relying on labels alone.
Q: Why do AI agent pipelines create new governance problems for identity teams?
A: Because agent pipelines often combine model calls, tool execution, and delegated access in one runtime path.
Q: What breaks when organisations rely on output filtering for AI governance?
A: Output filtering only limits what users see after data has already entered the system.
Practitioner guidance
- Implement pre-ingestion data classification Classify unstructured content before it is admitted to RAG pipelines, fine-tuning sets, or AI knowledge bases.
- Enforce data minimisation for AI inputs Remove or exclude records that are not necessary for the specific AI use case, especially where personal data, confidential business material, or regulated content can be avoided without reducing model utility.
- Review machine identities that move data Map the service accounts, tokens, and delegated permissions that can ingest content into AI systems.
What's in the full article
BigID's full article covers the operational detail this post intentionally leaves for the source:
- How its discovery workflow connects to more than 200 data sources without moving content out of place.
- The specific classification methods used to reduce false positives across large unstructured estates.
- Examples of access intelligence and remediation actions for documents, folders, and shared repositories.
- How policy enforcement is applied to data used in AI pipelines and retrieval systems.
👉 Read BigID's analysis of unstructured data governance for AI →
Unstructured data governance for AI: are your controls ready?
Explore further
Pre-ingestion control is the real AI governance boundary: once unstructured data enters a retrieval index or training set, the governance problem becomes structurally harder to unwind. BigID’s article reflects a wider pattern across AI programmes: organisations try to police outputs while ignoring the data estate that feeds the model. Practitioners should treat ingestion as the point where data security, IAM, and AI governance converge.
A question worth separating out:
Q: Who is accountable for governing data used in AI training and retrieval?
A: Accountability should sit with the business owner of the data, the security team responsible for classification and access, and the AI programme owner who decides what enters the pipeline. Frameworks such as GDPR and the EU AI Act push organisations toward demonstrable control over data minimisation, access, and lifecycle decisions.
👉 Read our full editorial: Governing unstructured data is now the AI risk control point