Unstructured data becomes a liability when organisations cannot tell what is inside it, who should access it, or whether it contains sensitive information. In that state, AI projects may amplify compliance risk, expose confidential content, and produce inaccurate outputs. Governance should advance before scale, especially when datasets are large, diverse, or weakly documented.
Why This Matters for Security Teams
Unstructured data becomes risky when AI systems can search, summarize, or infer across content that was never classified, scoped, or approved for machine use. In that state, the problem is not just storage sprawl. It is uncontrolled disclosure, weak provenance, and inconsistent access decisions that can surface sensitive material in prompts, embeddings, or generated outputs. NHI issues compound the problem because AI workloads rely on service accounts, tokens, and API keys to move through that data.
That is why guidance in the NIST AI Risk Management Framework matters here: organisations need to understand context before they automate decisions at scale. NHIMG research shows the same pattern in identity risk, where the Top 10 NHI Issues repeatedly include overexposed credentials and poor governance, not just isolated technical misconfigurations. In practice, many security teams encounter unstructured-data risk only after an AI pilot has already indexed material that should never have been machine-readable.
How It Works in Practice
The decision point is usually not whether unstructured data has value, but whether the organisation can make it safe enough for AI consumption. Current best practice is to treat unstructured repositories as high-risk until they are inventoried, classified, and tied to explicit access policies. That means knowing whether files contain secrets, regulated data, customer records, legal material, source code, or internal strategy. It also means understanding which NHIs can reach those stores, and whether those identities are over-privileged, stale, or shared.
Practical controls usually follow a staged pattern:
- Inventory the repositories, file shares, knowledge bases, and ticketing exports that the AI system can ingest.
- Classify the content by sensitivity, retention, and business owner before training, retrieval, or indexing begins.
- Restrict AI access with least privilege, strong workload identity, and short-lived credentials rather than broad service account access.
- Use policy checks at ingestion and retrieval time so sensitive content is blocked or redacted before it reaches the model.
- Log what the AI saw, what it returned, and which identity retrieved it, so investigations can trace exposure paths.
The operational logic is straightforward: AI increases value when it can retrieve relevant context, but it increases risk when it can retrieve everything else too. NHIMG’s 2024 ESG Report: Managing Non-Human Identities found that 72% of organisations have experienced or suspect a breach of non-human identities, which shows how quickly machine access turns into enterprise exposure. These controls tend to break down in large legacy content stores with no owner, no metadata, and no reliable way to distinguish sensitive material from routine documents.
Common Variations and Edge Cases
Tighter data controls often increase friction for search, analytics, and copilot-style workflows, requiring organisations to balance speed of AI adoption against the cost of curation and review. That tradeoff is especially visible in environments with mixed-quality content, because some repositories are worth enabling quickly while others should be excluded until governance improves.
There is no universal standard for this yet, but current guidance suggests treating these edge cases differently:
- High-value, low-sensitivity knowledge bases may be safe to expose with normal retrieval controls.
- Contract archives, HR files, incident tickets, and code repositories usually need stronger filtering, redaction, and owner approval.
- Data lakes and shared drives without clear lineage often create more risk than value because model output cannot be trusted to stay within intended boundaries.
- Unstructured data that contains secrets or credentials is especially dangerous because AI can surface material that should have been treated as an NHI issue first, not a content issue.
For governance teams, the important test is simple: if the organisation cannot explain what the data contains, who can access it, and why the AI needs it, the dataset is probably not ready. The NIST Cybersecurity Framework 2.0 and NHIMG’s OWASP NHI Top 10 both reinforce the same practical lesson: value only materialises when identity, access, and content governance move together.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Non-Human Identity Top 10, CSA MAESTRO and OWASP Agentic AI Top 10 address the attack and risk surface, while NIST AI RMF and NIST CSF 2.0 set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP Non-Human Identity Top 10 | NHI-01 | Unstructured data risk grows when machine identities overreach on access. |
| CSA MAESTRO | D1 | Agentic data use needs governance over retrieval, tools, and identity. |
| NIST AI RMF | AI risk management requires understanding data context before deployment. | |
| OWASP Agentic AI Top 10 | A03 | Agents can expose sensitive data through retrieval and prompt flows. |
| NIST CSF 2.0 | PR.DS | Data security controls map directly to unstructured content governance. |
Inventory NHI access to data stores and remove broad credentials before AI indexing begins.
Related resources from NHI Mgmt Group
Deepen Your Knowledge
Reviewed and updated by the NHIMG editorial team on August 28, 2026.
NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org