Use the least reversible option that still supports the business task. Redaction removes the data entirely, masking preserves readability with synthetic substitutes, and tokenization keeps reversibility inside a separate vault. The choice depends on whether the downstream workflow truly needs original values and whether the mapping store can be protected as a separate sensitive asset.
How the Three Techniques Differ in Practice
Redaction, masking, and tokenization all reduce exposure, but they do not do so in the same way. Redaction removes the value from the prompt or output entirely, which is the strongest privacy posture when the original value is not needed. Masking preserves enough structure for humans or downstream systems to read the record, while tokenization replaces the value with a surrogate that can be mapped back through a protected lookup service or vault.
The practical distinction is reversibility. Redaction is effectively one-way for the current workflow, masking is usually one-way for the reader but still reveals pattern and length, and tokenization is reversible only if the detokenization path remains protected. That makes tokenization closer to controlled substitution than true removal, so it belongs where business continuity requires the original value to be recoverable under policy.
For LLM privacy controls, the key question is not which method sounds most secure in the abstract, but which one still allows the model to do its job. If the task only needs classification, summarisation, or routing, redaction is often the cleanest option. If the task needs context but not the exact value, masking can preserve utility. If the workflow must later reidentify the original record, tokenization can preserve operational value without exposing the cleartext to the model.
Choosing the Least Reversible Option That Still Works
The strongest choice is the one that removes the most sensitive information without breaking the downstream use case. That usually means starting with redaction and only relaxing to masking or tokenization when a concrete business requirement proves that the original value, or a reversible surrogate, is needed. This is especially important for prompts, transcripts, logs, and retrieval corpora where LLMs can retain or regurgitate exposed content.
Redaction is best when the LLM does not need the value at all. It is also the safest default for free-text fields that may contain account numbers, passwords, API keys, health data, or other high-consequence content. The trade-off is obvious: once removed, the value cannot support downstream workflows unless another system already holds a clean copy.
Masking is useful when the model or reviewer needs shape, format, or partial context. A customer number, email domain, or last four digits may be enough for triage, matching, or human review. The downside is that partial disclosure can still leak identifying structure, and repeated masked records can enable correlation across sessions. Use it when recognisability has clear business value, not as a default compromise.
Tokenization is the right fit when the application needs a stable stand-in and a protected system can recover the original under strict control. Because the mapping store becomes a sensitive asset, tokenization shifts the risk rather than eliminating it. It works best when the token itself is useless outside the intended workflow and when the detokenization boundary is separate from the LLM runtime.
What Actually Breaks Privacy in an LLM Workflow
The failure mode is usually not the transform itself, but where the original data still exists and who can reach it. If redaction is performed too late, the model may already have seen the sensitive value. If masking is too weak, the LLM may still infer the underlying identity from surrounding context. If tokenization is reversible but the vault is overexposed, the system has only moved the problem to a different trust boundary.
Utility and privacy also pull against each other in retrieval, chat memory, and log retention. Data that is safe to show to a person may still be inappropriate to feed into a model if it can be memorised, summarised, or echoed back in another context. For that reason, privacy controls should be applied before ingestion wherever possible, not only at the output filter stage.
Risk and Threat Considerations
LLM privacy failures often come from over-sharing, weak separation between the model and source data, or reversible controls that are treated as if they were deletion. The main risk is that sensitive values remain recoverable either by the model, by an operator, or through a compromised detokenization path.
Failure mechanism: Sensitive data enters the prompt, context window, logs, or retrieval layer before being removed, or tokenization is deployed without sufficiently protecting the mapping store and access path.
Impact: The organisation can expose personal data, credentials, customer records, or confidential business content through model output, operator access, or later compromise of the token vault.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
NIST SP 800-53 Rev 5 and CIS Controls v8 set the technical controls, while ISO/IEC 27001:2022 defines the regulatory obligations.
| Framework | Control / Reference | Relevance |
|---|---|---|
| NIST SP 800-53 Rev 5 | SC-28 — Protection of Information at Rest | LLM privacy controls protect sensitive content stored in prompts, logs, and token vaults. |
| AC-6 — Least Privilege | Tokenization and detokenization paths should be limited to only the workflows that need recovery. | |
| IA-5 — Authenticator Management | LLM privacy programs often protect credentials and secret-like values that must be rotated or removed. | |
| Recommendation — Encrypt stored prompts, logs, and token mappings, then restrict access to the protected data stores. Restrict detokenization and vault access to the smallest set of approved services and operators. Rotate or revoke exposed secrets and prevent them from entering prompts, logs, or retrieval context. | ||
| ISO/IEC 27001:2022 | A.5.15 — Access control | Choosing redaction, masking, or tokenization depends on controlling who can recover sensitive values. |
| Recommendation — Define and enforce access rules for any system that can reverse a token or reveal original data. | ||
| CIS Controls v8 | CIS-3 — Data Protection | These techniques are data protection controls used to limit exposure of sensitive values in AI workflows. |
| Recommendation — Classify sensitive fields and apply the strongest transform that still preserves the required workflow. | ||
Practitioner Guidance
What to prioritise: Treat redaction as the default for any field the LLM does not truly need. Escalate to masking or tokenization only when you can state the downstream use case in one sentence and prove why less reversible handling would fail.
What to verify: Confirm that the privacy control runs before the content reaches the model, not just after generation, and that token mappings are stored separately from the LLM application, prompt logs, and general analytics systems. If a control depends on manual discipline, assume it will fail at scale.
Common mistake: Teams often use masking as a comfort layer while leaving enough structure for reidentification or correlation. If the value is high-risk, partial visibility is still a form of disclosure, not a safe compromise.
Practitioner takeaway: Use the least reversible control that preserves the business workflow, but treat any reversible mechanism as a protected asset with its own access rules, retention limits, and review path.
Related resources from NHI Mgmt Group
- Which compliance and security controls improve when organisations use data tokenization?
- What breaks when organisations rely on acceptable-use policies instead of technical controls for AI data privacy?
- What should organisations do when an LLM use case cannot be made fully privacy safe?
- When should organisations use NIST privacy framework-style controls instead of relying on informal judgment?
Deepen Your Knowledge
Free weekly newsletter
Subscribe to the NHI & AI Identity Journal
The latest on NHI and Agentic AI security – articles, research, breaches, news and events every week.
Bonus 33% off our NHI Course when you subscribe.
Reviewed and updated by the NHIMG editorial team on October 11, 2026.
NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org