Predictive coding is a machine learning approach that applies patterns from expert reviewer decisions to rank or classify additional documents. In legal review workflows, it helps identify likely relevant and nonrelevant material faster than manual review alone, especially when datasets contain millions of messages or files.
How predictive coding works
Predictive coding starts with a training set in which experienced reviewers label documents, then applies those patterns to score the rest of the corpus. The model does not replace legal judgment, it compresses the review problem by surfacing likely relevant material sooner.
Its value comes from prioritisation, not magic classification. In practice, predictive coding is usually one part of a broader review workflow that still depends on defensible document collection, review criteria, and quality checks on the resulting rankings.
Where predictive coding fits in legal review
Predictive coding is most useful when the corpus is too large for linear manual review to be efficient. It is commonly used in litigation, investigations, and other e-discovery style workflows where the goal is to identify responsive or privileged material at scale.
Because the approach learns from reviewer examples, the quality of the seed set and subsequent feedback loop matters. If the initial labels are inconsistent or the review protocol is poorly defined, the model can inherit those flaws and amplify them across the corpus.
Why predictive coding is distinct from simple keyword search
Keyword search matches terms; predictive coding infers relevance from patterns in the content and from prior reviewer decisions. That makes it better suited to issues where relevant documents use different vocabulary, mixed terminology, or indirect references.
This difference matters operationally. A keyword strategy can miss conceptually relevant items and can over-collect noisy results, while predictive coding can improve recall and reduce manual effort when the review objective is well defined and the training process is sound.
Quality, defensibility, and limitations
Predictive coding is only as defensible as the process around it. Teams need a clear review protocol, repeatable decision criteria, and validation steps that show the scoring process is working as intended for the matter at hand.
It also has practical limits. The method depends on reviewer expertise, representative training examples, and continued oversight, and it can be less reliable when the document population shifts materially or when the target issue is narrowly defined but sparsely represented.
Risk and Threat Considerations
Predictive coding creates process risk when legal teams treat the model score as a substitute for human review discipline. Poor seed sets, inconsistent labeling, or weak validation can drive false negatives that leave responsive documents undiscovered, or false positives that waste review time and inflate cost.
Failure mechanism: The model learns from the examples it is given, so biased or incomplete training data can mis-rank entire document classes and distort downstream review decisions.
Impact: Missed evidence, weaker defensibility, higher privilege or responsiveness risk, and longer remediation cycles if the review process has to be reopened.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP ASVS, NIST CSF 2.0 and NIST SP 800-53 Rev 5 set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP ASVS | V15 — Secure Coding and Architecture | Predictive coding is a model-driven review workflow that depends on sound system design and controls. |
| Recommendation — Design the review pipeline to preserve traceability, validation, and controlled change management. | ||
| NIST CSF 2.0 | GV.OV-01 — Outcomes and Oversight | The term hinges on governed oversight, validation, and accountability for model-assisted review decisions. |
| ID.RA-01 — Asset Vulnerabilities Are Identified and Recorded | Predictive coding risk rises when corpus characteristics and review weaknesses are not understood. | |
| PR.DS-01 — Data-at-Rest Is Protected | Document collections used for predictive coding often contain sensitive legal material that must be safeguarded. | |
| Recommendation — Establish oversight for how predictive coding outputs are validated and accepted in legal review. Identify corpus and workflow weaknesses that could distort training or ranking outcomes. Protect document repositories and review exports that feed predictive coding workflows. | ||
| NIST SP 800-53 Rev 5 | AU-6 — Audit Record Review, Analysis, and Reporting | Defensibility depends on reviewable records of what was trained, scored, and validated. |
| Recommendation — Retain and review logs that show how predictive coding decisions were produced and checked. | ||
Practitioner Guidance
Why practitioners should care: Predictive coding is a workflow control, not just a technology choice. Its real value depends on whether the review team can explain the training process, validate the output, and preserve a defensible record of how decisions were made.
Practitioner takeaway: Use it when scale makes manual review inefficient, but keep the protocol, sampling, and validation steps tight enough that the result can stand up to scrutiny.
Related resources from NHI Mgmt Group
Deepen Your Knowledge
Reviewed and updated by the NHIMG editorial team on September 27, 2026.
NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org