Join our Newsletter — 33% off our NHI Course

How should organisations start a PII compliance programme when they do not know where sensitive data is stored?

Start by locating PII across structured and unstructured systems, then map that data to the laws that actually apply to the organisation. A discovery-led approach gives teams the baseline needed for risk assessment, minimisation, safeguards, and response planning. Without that inventory, compliance work becomes partial, manual, and easy to miss.

Start with data discovery, not policy paperwork

The first practical step is to find where PII actually lives, across databases, file shares, ticketing exports, email archives, collaboration tools, and application logs. For most organisations, the biggest early failure is assuming the data inventory already exists in one system when it is really scattered across structured and unstructured stores, often with inconsistent labels and ownership.

Discovery works best when it is evidence-led. Combine automated scanning for known identifiers and sensitive patterns with sampling, business-owner interviews, and system-level review so you do not miss embedded PII in attachments, free text, or copied datasets. If you can only inventory the obvious repositories, the programme will undercount exposure and misstate the compliance baseline.

That baseline also needs to distinguish where PII is stored from where it is merely processed. A record may pass through logs, analytics pipelines, or support tools without those systems being the primary source of truth, but those copies still create compliance obligations. The programme starts to become useful when it can answer not just “where is the data?” but “which systems create, retain, duplicate, or export it?”

Once discovery produces a usable inventory, the next task is to map each PII category to the laws, contractual commitments, and internal handling rules that apply. That mapping is what turns a list of assets into a compliance programme, because the same dataset may be subject to different retention, disclosure, cross-border transfer, or security requirements depending on jurisdiction and business purpose.

A good mapping exercise identifies the minimum set of facts needed for action: data category, purpose, lawful basis or business justification, system owner, location, retention period, sharing routes, and escalation path. If those fields are missing, teams usually default to broad controls that are hard to maintain or to narrow controls that leave gaps. The point is to make the inventory operational, not merely descriptive.

When the legal scope is unclear, teams should treat that as a work item in the programme rather than a reason to delay it. Start with the jurisdictions and regulations that are clearly in play, then expand the mapping as discovery reveals additional processing locations or cross-border flows. This keeps the programme moving while preventing teams from building controls against the wrong standard.

Build control priorities from the inventory you can trust

After the first inventory and mapping pass, use it to prioritise the controls that reduce the most risk fastest: minimisation, retention cleanup, access restriction, encryption where appropriate, logging, and response playbooks for exposed data. This is where discovery becomes a governance tool, because teams can finally separate low-value data from high-impact data and stop applying the same treatment everywhere.

If you need a practical benchmark for why discovery matters, NHIMG’s Ultimate Guide to Non-Human Identities notes that only 5.7% of organisations have full visibility into their service accounts, a reminder that hidden assets are a common root cause of missed exposure. The same visibility problem appears in PII programmes when data is duplicated into tools nobody tracks well.

For programme design, the key decision is whether the inventory is accurate enough to support action. If not, keep the scope narrow and iterative: prove coverage on a few high-value systems, assign ownership, remediate obvious excess, then expand. A compliance programme that tries to cover every repository on day one often becomes a reporting exercise instead of a control programme.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

NIST CSF 2.0, CIS Controls v8 and NIST SP 800-63 set the technical controls, while ISO/IEC 42001:2023 and PCI DSS v4.0 define the regulatory obligations.

Framework Control / Reference Relevance
NIST CSF 2.0 ID.AM — Asset Management PII discovery depends on knowing where data assets reside and who owns them.
GV.OC — Organizational Context Mapping PII to applicable laws requires understanding business context and obligations.
PR.DS — Data Security Discovery feeds the controls used to protect sensitive data at rest and in transit.
Recommendation — Inventory data stores and owners before applying PII controls. Define the organisation’s regulatory context before setting compliance scope. Protect identified PII with handling, storage, and transfer controls that match sensitivity.
ISO/IEC 42001:2023 A.7 — Data and Information Governance The programme needs governed identification, handling, and retention of sensitive information.
Recommendation — Establish controlled data governance processes for discovery, classification, and retention.
CIS Controls v8 1 — Inventory and Control of Enterprise Assets A PII programme starts with discovering the systems that store or process sensitive data.
3 — Data Protection Once PII is found, protective controls should follow the sensitivity and exposure level.
Recommendation — Maintain a current inventory of systems that store or process PII. Apply data protection controls to identified PII repositories and exports.
NIST SP 800-63 IAL — Identity Proofing and Enrollment Assurance Levels When PII supports identity processes, assurance and collection scope affect handling obligations.
Recommendation — Limit collection to the minimum PII needed for the required assurance level.
PCI DSS v4.0 3 — Protect Stored Account Data Where payment-related PII or adjacent sensitive data is in scope, discovery is required before protection.
Recommendation — Locate and protect stored sensitive data before validating compliance.

Practitioner Guidance

What to prioritise: Start with systems that are both high-volume and high-risk, such as customer databases, shared file locations, support tooling, and log stores. Those are usually the places where undiscovered PII creates the largest compliance and response burden.

What to verify: Confirm that each discovered dataset has a named owner, a business purpose, and a retention rule. If any of those three are missing, the inventory is not yet reliable enough to support minimisation or deletion decisions.

What practitioners underestimate: Unstructured content often drives the largest blind spot. Searchable text, exports, attachments, and copied spreadsheets can carry more compliance risk than the system of record because they are easier to forget and harder to govern.

Practitioner takeaway: The fastest path to a usable PII compliance programme is not perfect legal analysis first, but a trusted inventory that is good enough to reveal where PII lives, who owns it, and which rules actually apply.