TL;DR: When sensitive attribute data is unavailable, organisations use proxy methods such as Bayesian Improved Surname Geocoding to estimate race or gender group membership, but the technique is intended for group-level disparity analysis and can mislead if treated as individual classification, according to Fiddler. The governance challenge is not only statistical accuracy but also accountability for how proxy inference is used in high-stakes decisions.
NHIMG editorial — based on content published by Fiddler: Identifying Bias When Sensitive Attribute Data is Unavailable: Techniques for Inferring Protected Characteristics
Questions worth separating out
Q: How should organisations use proxy methods for protected characteristics in fairness analysis?
A: Use them only for aggregate bias testing, not as a substitute for self-reported identity.
Q: Why do proxy methods create risk in regulated decision-making?
A: They can be accurate enough for population analysis while still being wrong for individuals.
Q: What do teams get wrong about inferring protected characteristics from available data?
A: They often confuse an estimate with evidence of identity.
Practitioner guidance
- Define approved proxy-use boundaries Write policy that limits inferred protected characteristics to fairness testing, research, or audit use cases, and explicitly forbids operational decisioning from proxy outputs.
- Document method confidence and limitations Require every proxy-based analysis to record the inference method, confidence assumptions, data sources, and known error modes before findings are used in governance reviews.
- Separate analysis from decision records Keep proxy-derived findings in an analytic layer that cannot be copied into case management, underwriting, hiring, or other production workflows without separate review.
What's in the full article
Fiddler's full blog covers the methodological detail this post intentionally leaves for the source:
- Bayesian Improved Surname Geocoding mechanics and the specific demographic inputs used to generate probability estimates
- Examples of how proxy methods are applied in lending and health care fairness assessments
- The article's cited research trail, including studies that question the accuracy and overstatement risk of proxy-based disparity analysis
- Context on how the method has been used in real enforcement and litigation settings
👉 Read Fiddler's analysis of inferring protected characteristics without sensitive data →
Inferring protected characteristics without sensitive data: what teams miss?
Explore further
Proxy inference is a governance instrument, not an identity truth source. Techniques like BISG exist because organisations need to test for disparate outcomes when sensitive attribute data is missing. That use case is legitimate, but only when the output is treated as an analytic approximation rather than a statement about who a person is. The practitioner conclusion is straightforward: separate fairness measurement from identity assertion.
A question worth separating out:
Q: Who should approve the use of inferred sensitive attributes in a fairness programme?
A: Privacy, legal, compliance, and fairness stakeholders should all review the method before it is adopted. The approval should cover whether the inference is necessary, whether the data inputs are appropriate, and whether the output will stay inside the intended analytic boundary.
👉 Read our full editorial: Inferring protected characteristics without sensitive data: bias risks