Privacy by design embeds data protection into the system or process from the start, so controls are considered during architecture and workflow decisions. Privacy by default applies the strongest privacy settings automatically, with minimal user effort. In practice, the first shapes how a system is built, while the second determines the baseline protections users get without manual tuning.
Why Privacy by Design and Privacy by Default Matter in AI and Data Governance
Privacy by design and privacy by default are often discussed together, but they solve different governance problems. Privacy by design is about shaping the architecture, data flows, retention rules, and access paths before a system goes live. Privacy by default is about the baseline posture users receive automatically, without needing to opt in or harden settings manually. In AI systems, that distinction matters because model training, prompt logging, telemetry, and downstream sharing can create privacy exposure long after deployment decisions are made.
For AI and data teams, the practical risk is that strong design intent can still fail if the default configuration leaks more data than necessary, or if the default settings are safe but the underlying workflow was never built with privacy constraints in mind. That is why privacy governance must cover both system architecture and operating defaults. The distinction is especially important where personal data, behavioural telemetry, or employee prompts may enter AI pipelines. Guidance from the NIST Cybersecurity Framework 2.0 helps teams anchor this in risk management rather than policy language alone.
NHIMG research shows how easily governance gaps become security gaps: in the 2024 ESG Report: Managing Non-Human Identities, 72% of organisations said they had experienced or suspected a breach of non-human identities. In practice, many teams discover privacy weaknesses only after a model, connector, or service account has already exposed more data than intended, rather than through deliberate privacy testing.
How Privacy by Design and Privacy by Default Work in Practice
Privacy by design is implemented during system design and engineering. It asks whether the AI use case needs the data at all, whether collection can be minimised, whether identifiers can be separated from content, and whether logs, prompts, and training corpora can be scoped to the smallest necessary dataset. For AI systems, this often means defining retention limits, masking sensitive fields before inference, controlling which workloads may access source data, and documenting lawful purpose and data lineage. The NIST Cybersecurity Framework 2.0 and NIST SP 800-53 Rev 5 Security and Privacy Controls both support this kind of control mapping.
- Build privacy requirements into architecture reviews, model selection, and vendor onboarding.
- Minimise collection and retention before data reaches model training or inference pipelines.
- Classify data so prompt logs, transcripts, and telemetry are handled differently from business content.
- Test whether integrations, exports, and fine-tuning paths preserve the intended privacy boundary.
Privacy by default operates after the design is in place. It means the safest reasonable settings are enabled automatically: restricted sharing, limited retention, opt-in rather than opt-out disclosure, and the lowest practical data exposure for each user or role. In regulated environments, default settings should also align with consent, notice, and access expectations under the EU General Data Protection Regulation (GDPR). For non-human identities that support AI workloads, default privilege should also be narrow, because over-permissioned service identities can bypass otherwise sound privacy design. Teams that neglect this often rely on manual hardening, which fails at scale when new agents, connectors, or environments inherit unsafe defaults. These controls tend to break down in fast-moving AI deployments where product teams clone environments faster than privacy settings are reviewed.
Common Variations and Edge Cases
Tighter privacy defaults often increase friction for analytics, product experimentation, and AI tuning, so organisations must balance data minimisation against operational usefulness. That tradeoff is real, and current guidance suggests there is no universal standard for every AI use case yet. The right answer depends on whether the system is customer-facing, internal, regulated, or using sensitive personal data.
One common edge case is model improvement through user interaction data. Privacy by design may allow collection for a legitimate purpose, but privacy by default should still keep sharing off unless the user or policy explicitly enables it. Another edge case is third-party AI services, where teams may control their own workflow but not the vendor’s retention or training settings. In those cases, privacy by design must include vendor due diligence and contract terms, not just internal engineering patterns. NHIMG’s Ultimate Guide to NHIs — Regulatory and Audit Perspectives is useful for connecting governance expectations to audit evidence, while the Ultimate Guide to NHIs — Key Research and Survey Results helps frame why weak defaults often persist even when policy looks complete.
The practical rule is simple: privacy by design prevents avoidable exposure, while privacy by default prevents avoidable misuse after deployment. Mature governance needs both.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Non-Human Identity Top 10 address the attack surface, NIST CSF 2.0, NIST SP 800-53 Rev 5 and NIST AI RMF set the technical controls, and EU AI Act define the regulatory obligations.
| Framework | Control / Reference | Relevance |
|---|---|---|
| NIST CSF 2.0 | GV.RM-01 | Privacy by design and default both depend on enterprise risk decisions. |
| NIST SP 800-53 Rev 5 | PT-2 | Privacy controls address data minimisation and purpose limitation. |
| NIST AI RMF | AI RMF covers governance of privacy risk in AI systems. | |
| OWASP Non-Human Identity Top 10 | NHI-01 | Non-human identities often carry excessive data access in AI pipelines. |
| EU AI Act | The AI Act increases scrutiny on data governance and transparency. |
Inventory service identities and reduce their access to only the data each workload needs.
Related resources from NHI Mgmt Group
- What is the difference between data access governance and DSPM in AI-enabled environments?
- What is the difference between disconnected privacy, security, and AI governance tools and a unified data command approach?
- What is the difference between private data storage and privacy by default in AI APIs?
- What is the difference between AI governance and AI runtime security?
Deepen Your Knowledge
Reviewed and updated by the NHIMG editorial team on September 1, 2026.
NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org