Start with data discovery and inventory. Organisations need to locate personal data across databases, file shares, email, cloud services, and third-party tools before they can assess legal basis, apply controls, or answer data subject requests. Once the data is mapped, teams can classify sensitivity, document processing activities, and prioritize remediation where regulated data is most exposed.
How to turn scattered LGPD data into a compliance map
The first practical step is to build a data inventory that is broad enough to catch where personal data actually lives, not just where it was intended to live. That means cloud databases, SaaS applications, shared drives, email, collaboration tools, and ad hoc repositories. For LGPD, the inventory is the foundation for lawful processing, retention, access control, and subject-rights response.
An inventory also creates the working boundary for the rest of the programme. Without it, teams tend to overfocus on policy documents while missing the systems that contain the highest-volume or highest-risk data, especially in business-owned SaaS and unstructured file stores.
A useful starting point is to define the minimum fields you need for each repository: system owner, data categories, likely personal-data types, geography, third-party access, retention expectations, and whether the repository can support search, export, deletion, and audit logging. That gives privacy, security, and application owners a shared structure for triage.
Why discovery must include cloud, SaaS, and unstructured content
LGPD obligations do not stop at the primary production database. In many organisations, the most difficult exposure sits in secondary copies, user-generated exports, mailbox attachments, shared folders, chat exports, and SaaS application data that sits outside central IT visibility. Discovery has to follow the data, not the platform boundary.
Cloud and SaaS environments also change the control problem. The organisation may not control the full stack, but it still needs to know what personal data is stored, who can access it, which processor or subprocessor is involved, and how deletion or export requests will be executed. For a cloud programme, a resource like Cloud Compliance Pulse 2025 is useful because it points practitioners back to cloud IAM hygiene and control visibility, which are often prerequisites for proving data handling discipline.
Unstructured repositories need extra attention because they are usually the least governed and the easiest to overlook. Email archives and document libraries often contain personal data in free text, attachments, or screenshots, so a discovery effort should include content scanning, naming-pattern review, and owner interviews rather than relying only on system catalogs.
What good looks like once the data map exists
Once data is mapped, the next job is not to produce a prettier spreadsheet, but to use the map to drive prioritisation. The highest-value outputs are: classifying sensitive records, identifying unnecessary copies, documenting processing activities, and separating business-critical repositories from legacy or orphaned stores. That is how discovery becomes a control decision rather than a one-time exercise.
Where the mapped data includes special-category or highly sensitive content, the organisation should connect the inventory to legal basis review, retention rules, and subject-rights workflows. LGPD programmes also benefit from a privacy-by-design view, because the inventory often reveals where minimisation, masking, or deletion should be applied before controls are extended.
For the legal side of that mapping, the GDPR guidance on data protection principles and DPIA thinking remains a useful comparator. The EU General Data Protection Regulation (GDPR) is not LGPD, but its structure reinforces the same practitioner logic: identify data first, then tie processing, controls, and impact assessment to the actual repositories. Where cloud services are part of the estate, the CSA Cloud Controls Matrix is a useful way to translate the inventory into concrete cloud control domains such as IAM, data security, and auditability.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
CSA Cloud Controls Matrix sets the technical controls, while GDPR defines the regulatory obligations.
| Framework | Control / Reference | Relevance |
|---|---|---|
| GDPR | Art.5 — Principles relating to processing of personal data | The answer centres on inventorying personal data before control application. |
| Art.25 — Data protection by design and by default | The answer emphasises discovery as the basis for privacy-by-design controls. | |
| Art.30 — Records of processing activities | The answer focuses on documenting where personal data sits across systems. | |
| Recommendation — Apply Art.5 principles to map data, minimise collection, and limit processing to defined purposes. Build privacy controls from the inventory so minimisation and default protections follow actual data flows. Maintain records of processing that reflect cloud, SaaS, and unstructured repositories. | ||
| CSA Cloud Controls Matrix | IAM — Identity and Access Management | Cloud and SaaS personal-data discovery depends on knowing access paths and owners. |
| DSP — Data Security and Privacy | The answer is about finding and classifying personal data across dispersed repositories. | |
| Recommendation — Use IAM controls to identify who can reach personal-data repositories and tighten access. Classify discovered personal data and apply handling controls based on sensitivity. | ||
Practitioner Guidance
What to prioritise: Start with the repositories most likely to combine high personal-data volume, broad user access, and weak owner visibility, because those are usually the fastest route to material LGPD exposure. Do not begin with every system equally; begin where discovery will change risk decisions quickly.
What to verify: Confirm that each repository can answer three questions, what personal data it holds, who can access it, and how it can be searched or removed. If a system owner cannot answer those questions, treat the repository as incomplete from a compliance standpoint until it is mapped.
Common mistake: Treating SaaS exports, shared mailboxes, and collaboration folders as operational noise rather than regulated storage. In practice, these are often where subject requests fail, retention breaks down, and sensitive copies multiply outside the main record system.
Practitioner takeaway: The right first milestone is not policy approval, it is a usable inventory that makes LGPD obligations executable across the real data estate, including the places where the business created shadow copies.
Related resources from NHI Mgmt Group
- How should security teams implement GDPR compliance when personal data is spread across SaaS, cloud, and AI tools?
- How should organisations start building a GDPR compliance programme when data flows are spread across teams and systems?
- How should SaaS teams implement DPDP compliance when they process personal data across cloud and GenAI systems?
- How should security teams prioritize data discovery for CCPA compliance when personal information is spread across cloud and on-prem systems?