Use the score as a decision control, not as a truth statement. Set thresholds against your own labeled data, route borderline cases to review, and keep deterministic rules for exact identifiers such as matching email addresses. In messy directories, a calibrated probability is most useful when it can separate auto-accept, human review, and reject paths in a repeatable way.
Using a calibrated score as a control, not a verdict
A calibrated identity-matching score is most useful when it drives a repeatable action, not when it is treated as proof of sameness. In large directories, that means the score should separate auto-accept, review, and reject paths, while exact identifiers such as an email address or employee ID stay governed by deterministic rules. The practical goal is to reduce false merges without turning every mismatch into a manual case.
Calibration matters because the same raw score can mean different things across datasets, tenants, or matching rules. Security teams should define the score as “estimated probability of correct match” only after validating it against local labeled data. That keeps the threshold anchored to observed error rates rather than vendor language or intuition.
How to set thresholds without over-merging
Thresholds should be chosen from your own false-positive tolerance, not from the score distribution alone. The safest pattern is to reserve a high-confidence band for automatic merges, a low-confidence band for rejection, and a middle band for analyst review. That middle band is where most false merges are prevented, because it catches plausible but incomplete matches before they become identity corruption.
For directories with messy attributes, use exact-match fields as hard gates before probabilistic logic. If two records disagree on a strong identifier, the score should not override that conflict unless there is an explicit exception rule and a documented reviewer path. This keeps the model from “winning” against better evidence.
- Set thresholds from labeled examples that resemble your production data.
- Use separate bands for auto-accept, review, and reject.
- Treat strong identifier conflicts as blockers unless a human explicitly approves an exception.
- Revisit thresholds after directory growth, acquisition, or source-system changes.
What false merges usually look like in practice
False merges often happen when the model overweights partial similarity, such as name variants, shared domains, stale titles, or recycled addresses. The operational problem is not only a wrong record, but a wrong trust decision: permissions, notifications, audit trails, and downstream workflow actions can all attach to the wrong person or account.
That is why teams should inspect merge candidates for failure modes that are common in large directories: duplicate names across business units, shared mailboxes, contractors with short tenure, and records imported from multiple systems with inconsistent normalization. A calibrated score helps, but only if the surrounding rules recognise that identity data quality is uneven by source.
When merging affects access, privileges, or recovery paths, NIST Cybersecurity Framework 2.0 and NIST SP 800-53 Rev 5 Security and Privacy Controls are useful anchors for separating governance, access control, and auditability expectations from the scoring logic itself.
Risk and Threat Considerations
False merges can create a high-impact identity collision: access, approvals, alerting, or recovery actions may be attributed to the wrong record, and the mistake can spread across connected systems. The main risk is not just bad data quality, but mistaken trust, because one merged profile can inherit privileges or visibility it should never have had.
Failure mechanism: A calibration score that is used as an automatic truth signal will sometimes overrule conflicting evidence, especially in noisy directories with partial overlaps, stale attributes, or recycled identifiers. That creates accidental account binding and can hide a bad merge until an access event, audit review, or incident exposes it.
Impact: The wrong person can receive access, notifications, attestations, or rollback paths, while the right one can lose continuity of history and accountability. At scale, repeated false merges also poison downstream matching, because bad labels become training or review precedent for future decisions.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
NIST CSF 2.0 and NIST SP 800-53 Rev 5 set the technical controls, while ISO/IEC 27001:2022 defines the regulatory obligations.
| Framework | Control / Reference | Relevance |
|---|---|---|
| NIST CSF 2.0 | GV.OC-01 — Organizational Context | Identity matching thresholds depend on local directory context and business tolerance. |
| Recommendation — Define acceptable match-error thresholds based on your directory's business use cases. | ||
| NIST SP 800-53 Rev 5 | IA-5 — Authenticator Management | Directory merges can affect credentialed identity records and lifecycle evidence. |
| AU-2 — Event Logging | Merge decisions need auditability to detect and investigate false identity joins. | |
| Recommendation — Require authoritative identity evidence before merging records that support access. Log merge decisions, reviewers, and overrides for later audit and rollback. | ||
| ISO/IEC 27001:2022 | A.5.15 — Access control | Identity merges can alter who receives or retains access, so controls must be governed. |
| A.8.15 — Logging | Probabilistic merges need traceable decision records for investigation and assurance. | |
| Recommendation — Enforce access-impacting merge approvals through documented control procedures. Record merge inputs, thresholds, and overrides for traceability. | ||
Practitioner Guidance
What to prioritise: Decide first which fields are authoritative and which are merely supportive. A calibrated score should only arbitrate among records that already pass your hard identity checks, otherwise the automation will quietly turn ambiguity into false certainty.
What to measure: Track false-merge rate, review overturn rate, and the share of merges that depended on borderline scores. If the review queue is small but the overturn rate is high, your threshold is probably too permissive or your labeled set is not representative.
Decision rule: If a merge would affect entitlements, recovery, or audit attribution, require a stronger evidence set than the score alone, even when the probability looks high. If the merge only improves deduplication and cannot change authority or access, you can tolerate a lower-risk review path.
Practitioner takeaway: The safest directory program treats probabilistic matching as a triage mechanism, then uses hard identifiers and human review to prevent a confident but incorrect merge from becoming a durable truth.
Related resources from NHI Mgmt Group
- How should security teams use machine learning without creating too many false declines?
- How should security teams use AI-driven risk decisioning without creating too many false positives for trusted users?
- How should security teams use regular expressions to discover sensitive data without creating too many false positives?
- How should security teams use behavioral analysis and AI in email security without creating too many false positives?
Deepen Your Knowledge
Reviewed and updated by the NHIMG editorial team on September 30, 2026.
NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org