TL;DR: Automated data discovery and classification were used before workloads moved to AWS in a financial services migration, reducing blind spots across sensitive data and embedded secrets while feeding findings into Security Hub, according to BigID. The core lesson is that cloud modernisation fails when governance starts after migration rather than before.
At a glance
What this is: This is a case study showing that pre-migration data discovery and classification can reduce cloud risk by identifying sensitive data and secrets before workloads move.
Why it matters: It matters to IAM practitioners because data visibility, secrets governance, and access control all depend on knowing what exists before permissions, workloads, and controls are translated into cloud.
By the numbers:
- Only 44% of organisations have implemented any policies to manage their AI agents, despite 92% agreeing that governing AI agents is critical to enterprise security.
👉 Read BigID's analysis of visibility-first cloud migration and data governance
Context
Cloud migration exposes a basic governance problem: organisations often move data and workloads before they have a defensible inventory of what is actually inside the estate. In finance, that gap is especially risky because sensitive data, embedded secrets, and access pathways can spread across structured and unstructured systems long before the cloud programme begins.
The article is about using data discovery and classification to change migration order, not just migration tooling. That matters to identity and access teams because secret discovery, entitlement scope, and data sensitivity all shape how permissions should be designed before workloads enter AWS.
The starting position described here is common, not exceptional. Many enterprises discover that their cloud risk is less about the destination platform and more about the visibility deficit they carried into it.
Key questions
Q: What breaks when cloud migration starts before data discovery?
A: Cloud migration starts to fail when teams move applications before they know what sensitive data, secrets, and regulated records already exist. The result is inherited risk, weak prioritisation, and access controls built on incomplete assumptions. A defensible migration plan needs inventory and classification first, then workload movement and control translation.
Q: Why do embedded secrets increase cloud migration risk?
A: Embedded secrets matter because they are credentials, not just data. If API keys, tokens, or certificates are hidden in files, code, or storage, they can create immediate access paths once the environment is copied into cloud services. Discovery is essential because you cannot rotate or revoke what you have not found.
Q: How do you know if classification is actually reducing cloud risk?
A: Classification is working when it changes operational decisions, not when it just produces labels. Look for fewer unidentified sensitive assets, faster remediation of exposed secrets, and migration gates that stop risky workloads until ownership and sensitivity are clear. If findings do not alter prioritisation, the programme is decorative rather than controlled.
Q: How should security teams maintain identity assurance during cloud migration?
A: Security teams should treat migration as an identity control redesign, not just a platform move. Preserve strong MFA, recovery assurance, and admin access rules across both cloud and on-premises environments, then verify that fallback paths do not reduce assurance below the primary sign-in standard. Continuity matters more than cloud placement.
Technical breakdown
Data discovery before migration
Data discovery is the process of locating and identifying sensitive information across legacy and cloud estates before workloads are moved. In this case, the key mechanism is classification first, migration second. That changes the security model because controls can be aligned to actual data sensitivity rather than assumed structure. Without discovery, cloud migration often copies ambiguity into a new environment, which makes governance reactive and incomplete. Practical implication: build a data inventory before replatforming so the migration plan is driven by sensitivity, not just application dependency.
Practical implication: Build a data inventory before replatforming so the migration plan is driven by sensitivity, not just application dependency.
Embedded secrets and exposed access paths
Secrets are credentials, tokens, API keys, and certificates that can create direct access if exposed in storage or application layers. In cloud environments, the main risk is that secrets are often hidden inside data stores, code, or unstructured repositories, where they are easy to overlook during migration. Automated classification reduces that blind spot by surfacing where secrets exist and which assets are high risk. Practical implication: treat secret discovery as a migration dependency, not a post-move cleanup task.
Practical implication: Treat secret discovery as a migration dependency, not a post-move cleanup task.
Security Hub as a shared control plane
When findings are pushed into AWS Security Hub, the goal is to centralise risk signals so cloud and data security teams can act from a common view. This is not just alert aggregation. It is a governance pattern that connects data sensitivity, posture, and remediation into a single operational loop. In identity terms, this is where visibility influences access decisions, because sensitive data should drive who and what can touch it. Practical implication: route classified findings into operational tooling that can trigger consistent remediation and access review.
Practical implication: Route classified findings into operational tooling that can trigger consistent remediation and access review.
Threat narrative
Attacker objective: The attacker objective is to find hidden credentials or sensitive data that grant access to cloud systems and increase the value of any compromise.
- Entry occurs when legacy data stores and unstructured repositories contain embedded secrets or sensitive data that were never fully inventoried.
- Escalation happens when that hidden data is moved into cloud services without classification, expanding the reach of any exposed credential or mislabelled asset.
- Impact is a larger cloud blast radius, because sensitive records and access paths are now distributed across the new environment with weaker governance than intended.
NHI Mgmt Group analysis
Data visibility is now a prerequisite for cloud governance: moving workloads before understanding the data estate simply relocates risk into a more scalable environment. Classification gives security teams a basis for control selection, retention decisions, and access scoping. For IAM and cloud teams, the practitioner conclusion is clear: cloud migration should begin with inventory, not infrastructure.
Embedded secrets create an identity problem before they become a cloud problem: once API keys, tokens, and credentials are scattered across storage and application layers, they function like unmanaged non-human identities. That is why data security and identity governance converge in migration programmes. If the organisation cannot find the secret, it cannot govern the access it enables.
Continuous classification is the named control gap this case exposes: the failure mode is not simply poor migration planning, but an absence of continuous data awareness across changing cloud estates. Data sensitivity changes what should be protected, logged, and reviewed. Practitioners should treat ongoing discovery as a governance control, not a one-time project output.
Security consolidation only helps when it reduces blind spots, not when it hides them: fewer tools can improve response only if the remaining control plane preserves both data context and operational reach. The value is in prioritising risk by sensitivity, not just by infrastructure misconfiguration. The practitioner conclusion is to test whether consolidation improves decision quality before it changes tooling count.
What this signals
Continuous data discovery is becoming a governance control, not a migration nicety: once cloud programmes start moving regulated data, the security question shifts from where the data lives to whether the organisation can still explain what it contains and who can reach it. That is where identity and data governance meet, especially around secrets and service access.
Hidden credentials behave like unmanaged non-human identities: the moment embedded secrets are discovered in storage or application layers, they should be handled as lifecycle-bearing access assets. That is why the lifecycle view in the NHI Lifecycle Management Guide matters even in a cloud migration story.
The operational signal to watch is whether data classification changes prioritisation in the cloud programme. If labelled data still moves without ownership, review, and remediation gates, the organisation has visibility without control, and the risk merely changes shape instead of shrinking.
For practitioners
- Inventory sensitive data before workload migration Map PII, personal data, and application secrets across structured and unstructured repositories before moving any workloads into AWS or another cloud platform.
- Classify secrets as governed access-bearing assets Treat embedded API keys, tokens, and certificates as access paths that require discovery, ownership, and lifecycle handling alongside the data they protect.
- Feed classified findings into operational response tools Push discovery results into a shared security console such as Security Hub so cloud and data teams work from one prioritized view of sensitive exposure.
- Tie migration gates to sensitivity thresholds Block or delay migration steps when critical data classes remain unidentified, unlabeled, or outside approved handling patterns.
Key takeaways
- This case shows that cloud migration without data visibility simply scales existing risk instead of reducing it.
- Embedded secrets and unclassified sensitive data create an identity and access problem before they become a platform problem.
- Practitioners should make discovery, classification, and remediation gates part of the migration path, not a post-move cleanup exercise.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
NIST CSF 2.0, NIST SP 800-53 Rev 5 and CIS Controls v8 set the technical controls, while ISO/IEC 27001:2022 define the regulatory obligations.
| Framework | Control / Reference | Relevance |
|---|---|---|
| NIST CSF 2.0 | PR.AC-4 | Access control decisions should reflect classified data sensitivity during migration. |
| NIST SP 800-53 Rev 5 | AC-6 | Least privilege is central when sensitive data and secrets are spread across cloud workloads. |
| ISO/IEC 27001:2022 | A.8.2 | Information classification is directly relevant to this migration-first governance model. |
| CIS Controls v8 | CIS-3 , Data Protection | Data protection controls align with discovery of sensitive records and secrets before migration. |
Apply CIS-3 to inventory, classify, and protect sensitive data before replatforming.
Key terms
- Data Discovery: Data discovery is the process of finding where information lives across cloud, SaaS, endpoints, backups, and analytics systems. In practice, it creates the inventory that makes classification, access decisions, recovery planning, and AI governance possible rather than speculative.
- Embedded Secret: An embedded secret is a credential, token, API key, or certificate that has been placed inside an image, file, or configuration bundle. Once exposed, it can outlive the original control that created it, so revocation and rebuilds become necessary even if the underlying system is later patched.
- Data classification: Data classification is the process of labelling information according to sensitivity, regulatory impact, or business value so controls can be applied consistently. For AI governance, it allows policy to follow the data into prompts, sessions, and destinations rather than relying on brittle text matching.
- Security Hub: Security Hub is the operational consolidation point where findings from multiple controls can be reviewed together. In this context, it matters because it connects data risk signals to cloud posture actions, helping teams prioritise response using a shared security view.
What's in the full article
BigID's full article covers the operational detail this post intentionally leaves for the source:
- How the discovery and classification workflow was applied across legacy systems and AWS services before migration
- How findings were operationalised through AWS Security Hub for triage and response
- How the programme reduced blind spots across structured databases and unstructured storage
- How the organisation aligned data sensitivity with cloud remediation priorities
Deepen your knowledge
The NHI Foundation Level course, the industry's only accredited NHI security programme, covers NHI governance, identity lifecycle, and secrets management for practitioners who need to connect access control with operational reality. It helps security teams translate identity principles into controls that hold across cloud, application, and service accounts.
Published by the NHIMG editorial team on August 20, 2026.
NHI Mgmt Group — the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org