Test the model on combinations of protected attributes, not only on one attribute at a time. The goal is to expose subgroup harms that disappear in aggregate reporting. Use the same decision threshold, compare the same outcome measure across intersections, and retain the evidence so governance teams can review the pattern, not just the headline metric.
Why This Matters for Security Teams
When protected attributes overlap, a model can look acceptable overall while failing specific populations in ways that are easy to miss in a single-metric review. That is why fairness testing needs to move beyond aggregate accuracy and into subgroup analysis that reflects how real people are affected. Current guidance from the NIST Cybersecurity Framework 2.0 reinforces the broader governance principle: controls only matter if they are measurable, repeatable, and tied to accountability.
For AI systems, the practical risk is not just discrimination claims. It is also avoidable operational error, inconsistent decisioning, and loss of trust in automated workflows. Teams often test one protected attribute at a time because it is simpler to report and easier to explain, but that approach can hide compounded harms. A model may appear balanced across gender and across age, yet still underperform for older women, disabled applicants from a minority ethnic group, or other intersecting subgroups.
Security, legal, and model-risk functions should treat intersectional testing as a governance requirement, not an optional refinement. If the organisation cannot show how fairness was evaluated across overlapping attributes, it cannot credibly claim that its testing reflects actual deployment conditions. In practice, many teams encounter intersectional bias only after complaints, appeals, or adverse outcome reviews have already exposed the gap, rather than through intentional pre-release testing.
How It Works in Practice
Effective fairness testing starts by defining the attributes that are lawful, relevant, and sufficiently represented in the evaluation data. Teams then create intersectional slices, such as combinations of sex, age band, disability status, ethnicity, or geography, and apply the same decision threshold across all slices so results remain comparable. The key is consistency: if thresholds, label definitions, or input windows change between groups, the comparison loses value.
Practitioners should measure more than one metric. Depending on the use case, that may include selection rate, false positive and false negative rates, calibration, or error disparity across intersecting groups. Where possible, the testing set should reflect operational reality, not a sanitized sample that removes difficult cases. Evidence should be retained in a form that supports audit, challenge, and re-testing when the model changes.
- Test single attributes and intersections, then compare the deltas, not only the headline result.
- Use the same thresholding logic for all slices unless there is a documented, justified reason not to.
- Record the sample size for each intersection so small groups are not overinterpreted.
- Check whether data quality, proxy variables, or missing labels are driving the apparent disparity.
- Link findings to governance actions, including model approval, remediation, and periodic revalidation.
This approach aligns with the AI risk management principles in NIST Cybersecurity Framework 2.0 when fairness is treated as part of accountable control design rather than an isolated ethics review. It also fits the testing mindset used in AI governance programs that track evidence, exceptions, and remediation over time. These controls tend to break down when intersectional groups are too small, labels are incomplete, or the deployed population shifts faster than the evaluation dataset.
Common Variations and Edge Cases
Tighter fairness testing often increases review effort, data preparation cost, and the chance of inconclusive results, so organisations have to balance stronger assurance against statistical uncertainty. That tradeoff is especially visible when several protected attributes intersect and some slices contain only a handful of cases.
There is no universal standard for the exact number of intersections that must be tested, so current guidance suggests prioritising combinations that are legally relevant, operationally material, and reasonably supported by data. Best practice is evolving for synthetic data, reweighting, and threshold adjustment, and those methods should be clearly labelled if they are used.
Some environments need extra caution. In high-impact decisions, such as hiring, lending, admissions, or fraud screening, a single unfair subgroup result may justify blocking deployment until remediation is complete. In lower-risk internal tools, the team may accept limited residual disparity if it is documented, monitored, and tied to a formal risk decision. Where human review is available, teams should still test whether reviewers inherit the same subgroup bias as the model. The strongest programs treat fairness findings as living evidence, not a one-time certification.
For governance alignment, the most useful next step is to connect intersectional testing results to model approval criteria, monitoring triggers, and escalation paths so that unusual subgroup outcomes lead to action rather than a stored report.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
MITRE ATLAS and OWASP Agentic AI Top 10 address the attack surface, NIST AI RMF and NIST AI 600-1 set the technical controls, and EU AI Act define the regulatory obligations.
| Framework | Control / Reference | Relevance |
|---|---|---|
| NIST AI RMF | AI RMF addresses governance and measurement of model impacts across affected groups. | |
| MITRE ATLAS | ATLAS helps teams think about adversarial manipulation that can distort fairness evaluation data. | |
| OWASP Agentic AI Top 10 | Agentic systems can amplify biased decisions when tools and actions inherit unfair model outputs. | |
| NIST AI 600-1 | The GenAI profile supports testing and monitoring for harmful or biased model behaviour. | |
| EU AI Act | The EU AI Act raises governance expectations for high-risk systems affecting protected groups. |
Validate GenAI outputs across intersecting groups before release and during ongoing monitoring.
Related resources from NHI Mgmt Group
- How should security teams test AI agents that can call tools and APIs?
- How should security teams govern workload identity federation across multiple AI APIs?
- How should security teams govern AI agents that use multiple identity layers?
- How should security teams govern AI agents that can invoke multiple tools in one session?
Deepen Your Knowledge
Reviewed and updated by the NHIMG editorial team on August 20, 2026.
NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org