The measurement phase of the NIST AI Risk Management Framework, focused on evaluating whether an AI system is fair, accurate, explainable, and secure in practice. It relies on metrics, benchmarks, and assessments to surface risk signals and verify that the system behaves as intended over time.
Expanded Definition
Measure Function is the part of the NIST Cybersecurity Framework 2.0-aligned AI governance cycle that turns broad assurance goals into observable evidence. In practice, it asks whether an AI system is performing as intended against defined criteria, rather than assuming that training-time validation is enough. For NHI Management Group, the important distinction is that measurement is not a one-off test: it is an ongoing discipline that combines metrics, tests, benchmarks, red-team style probing, and human review to detect drift, bias, weak explainability, and security exposure.
Definitions vary across vendors on where measurement ends and monitoring begins, but the core idea is stable: measurement produces risk signals that decision-makers can act on. It sits between policy intent and operational response, giving governance teams a way to compare actual system behavior with acceptable thresholds. In AI security contexts, this often includes performance regression checks, prompt or output quality evaluation, and safety assessments tied to specific use cases. The most common misapplication is treating model accuracy as a complete proxy for Measure Function, which occurs when organisations ignore fairness, robustness, and security indicators after initial validation.
Examples and Use Cases
Implementing Measure Function rigorously often introduces review overhead and tooling complexity, requiring organisations to weigh faster deployment against stronger assurance.
- An enterprise measures an LLM used for customer support against hallucination rate, refusal quality, and sensitive-data leakage, then revisits thresholds after production drift is observed. This is consistent with the measurement intent described in NIST CSF 2.0 style governance.
- A bank evaluates an AI triage model for bias across protected groups, using repeated test sets to confirm that service outcomes remain stable when input patterns change.
- A security team runs adversarial prompt tests against an internal assistant to check whether tool-use boundaries, policy filters, and data handling controls still hold after model updates.
- A regulated organisation benchmarks explainability outputs so reviewers can trace why high-impact recommendations were produced, especially where human override is required.
- An MLOps team compares pre-release and post-release metrics to detect whether a retrained model has introduced regressions in accuracy, latency, or unsafe output patterns.
Why It Matters for Security Teams
Measure Function matters because security teams need evidence, not assumptions, when AI systems influence decisions, expose data, or trigger downstream automation. Without measurement, risk management becomes reactive and blind to model drift, emerging prompt abuse, broken guardrails, and changes in user behaviour that alter the threat surface. It also supports governance decisions by showing whether controls are actually working under realistic conditions, not just in a lab. This is where AI assurance intersects with broader cyber governance: measurement provides the operational signal that feeds incident response, change management, and executive oversight.
For identity-adjacent deployments, measurement becomes especially important when AI agents use tools, secrets, or privileged workflows, because small failures can create large authorization or data-handling consequences. Teams should think of it as the bridge between design-time intent and runtime confidence, particularly where NIST Cybersecurity Framework 2.0 functions are being adapted to AI governance. Organisations typically encounter uncontrolled outputs, audit gaps, or trust breakdowns only after a harmful decision or public incident, at which point Measure Function becomes operationally unavoidable to address.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
NIST AI RMF, NIST AI 600-1, NIST CSF 2.0, NIST SP 800-53 Rev 5 and NIST SP 800-63 set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| NIST AI RMF | Measure is a core function of the NIST AI Risk Management Framework. | |
| NIST AI 600-1 | The GenAI profile extends AI governance measurement expectations for generative systems. | |
| NIST CSF 2.0 | GV.OC, DE.CM | CSF governance and monitoring concepts support ongoing measurement of security outcomes. |
| NIST SP 800-53 Rev 5 | CA-7 | Continuous monitoring control supports evidence-based assessment of control effectiveness. |
| NIST SP 800-63 | Identity assurance guidance is relevant where AI measurement evaluates verification or authentication flows. |
Measure identity-related AI decisions against assurance thresholds and review exceptions promptly.
Related resources from NHI Mgmt Group
- How should security teams measure the business value of identity security?
- How should organisations measure identity security ROI beyond license savings?
- How should security teams measure AI success without creating blind spots?
- How should security teams measure whether AI is helping rather than hiding risk?
Deepen Your Knowledge
Reviewed and updated by the NHIMG editorial team on August 24, 2026.
NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org