TL;DR: NIST’s evaluation of a facial age estimation model found a Mean Absolute Error improvement from 3.102 to 2.615 on the mugshot dataset, with the model moving from 12th to 3rd place and narrowing the gap to first place to 0.2 years, according to Yoti. The result underscores how independent testing, bias analysis, and demographic performance matter as much as headline accuracy in regulated identity checks.
NHIMG editorial — based on content published by Yoti: latest evaluation of facial age estimation by NIST
By the numbers:
- Mean Absolute Error has improved from 3.102 to 2.615 for the NIST mugshot data set.
- NIST evaluates models across over 20 million images.
Questions worth separating out
Q: How should organisations use facial age estimation in regulated identity workflows?
A: Use it as one control in a layered assurance process, not as the only decision maker.
Q: Why does independent testing matter for biometric age checks?
A: Independent testing shows how the model behaves outside the vendor’s own lab conditions.
Q: What do security and identity teams get wrong about age verification?
A: They often treat it as a one-time onboarding check instead of an ongoing governance process with evidence, testing, and jurisdiction-specific rules.
Practitioner guidance
- Set acceptance thresholds for age estimation outputs Define separate approval thresholds for overall error, subgroup variance, and scenario robustness before any production rollout.
- Require independent benchmark evidence Demand external evaluation results for the exact model version and keep the dataset context with the approval record.
- Test demographic and scenario variance Review performance across age bands, gender proxies, image quality, and presentation changes such as glasses or expressions.
What's in the full article
Yoti's full article covers the underlying evaluation detail this post intentionally leaves for the source:
- The model-by-model comparison tables for MAE across multiple NIST datasets.
- The demographic breakdowns by age group, gender proxy, and skintone proxy that inform the bias discussion.
- The robustness testing results for expression changes, glasses, and other presentation variations.
- The longer explanation of how the model training strategy changed over time and why the vendor says it improved.
👉 Read Yoti's evaluation of facial age estimation performance in NIST testing →
NIST facial age estimation results: what do identity teams do now?
Explore further
Independent evaluation is now the minimum acceptable trust signal for age-based identity controls. When a facial age estimation model is used to support access or compliance decisions, vendor-led testing is not enough. External benchmarking exposes whether performance survives different datasets, image conditions, and user populations. Practitioners should treat third-party evaluation as a baseline requirement, not a differentiator.
A few things that frame the scale:
- 80% of identity breaches involved compromised non-human identities such as service accounts and API keys, according to the Ultimate Guide to NHIs.
- 91.6% of secrets remain valid five days after the targeted organisation is notified, according to Ultimate Guide to NHIs , Key Research and Survey Results.
A question worth separating out:
Q: Who should approve the use of facial age estimation for access decisions?
A: Approval should sit with the identity, risk, and compliance owners together, not only the product team. The decision should cover acceptable error ranges, demographic testing, review cadence, and what happens when the model falls outside tolerance. That makes accountability explicit before the control goes live.
👉 Read our full editorial: NIST age estimation evaluation raises the bar for face checks