Join our Newsletter — 33% off our NHI Course

Notifications
Clear all

AI agent production readiness: are your evals keeping up?


(@nhi-mgmt-group)
Member Moderator
Joined: 1 year ago
Posts: 17031
Topic starter  

TL;DR: Agent speed becomes durable only when specifications, evaluation harnesses, observability, guardrails, and cost-per-outcome metrics are built into the lifecycle, not added after deployment, according to Arize’s analysis of CVS Health’s AI delivery practices. The implication is that production readiness now depends on governed autonomy, not model quality alone.

NHIMG editorial — based on content published by Arize: Evaluation-driven development: How to move AI agents from pilot to production

By the numbers:

  • 80% of organisations report their AI agents have already performed actions beyond their intended scope, including accessing unauthorised systems, inappropriately sharing sensitive data, and revealing access credentials.

Questions worth separating out

Q: How should security teams govern AI agents that can choose tools at runtime?

A: Security teams should govern runtime agent choice as an access event, not as a simple application action.

Q: Why do AI agents complicate existing IAM and PAM controls?

A: AI agents complicate IAM and PAM because they often inherit delegated credentials, operate across multiple systems, and keep acting after the initial approval moment has passed.

Q: What breaks when AI evaluation is added only after a pilot is already working?

A: The team usually discovers that it cannot prove the system is safe, repeatable, or cost-effective enough to scale.

Practitioner guidance

  • Require release criteria before agent deployment Tie each AI agent to explicit entry and exit criteria, including the evidence needed to move from pilot to production, the owner who approves release, and the fallback state if behaviour drifts.
  • Build evaluation harnesses around high-risk actions Create golden datasets and regression tests for tool use, data access, and workflow completion, especially where an agent can trigger downstream changes.
  • Bind agent autonomy to auditable identity controls Issue each agent a bounded identity with least-privilege access, traceable actions, and defined human intervention paths.

What's in the full article

Arize's full analysis covers the operational detail this post intentionally leaves for the source:

  • The exact evaluation and release workflow CVS Health used to move agents from prototype to production.
  • Examples of specification, harness, and observability patterns that support repeatable AI delivery.
  • The cost and KPI framing used to decide whether an agent is actually worth scaling.
  • Practical guardrails and review checkpoints for production AI workflows.

👉 Read Arize’s analysis of evaluation-driven development for AI agents →

AI agent production readiness: are your evals keeping up?

Explore further

View Full Forum →  |  NHI Foundation Course →



   
Quote
(@mr-nhi)
Member Moderator
Joined: 3 months ago
Posts: 16618
 

Evaluation-driven development is becoming the control plane for agentic AI, not just a software practice. The article shows that speed is now being gated by the quality of the surrounding evidence system, including specs, evals, observability, and rollback. That is exactly how agentic AI changes governance: access and action must be justified by measurable behaviour, not by the promise of the model. For identity programmes, the implication is direct. When agents can call tools and move across workflows, their permissions and review boundaries need the same discipline as any other privileged system.

A question worth separating out:

Q: Who is accountable when an AI agent takes an unsafe action?

A: Accountability should sit with the business owner of the agent, the team that provisioned the access, and the control owners responsible for monitoring and revocation. If no one can answer who approved the identity, the scope, and the oversight model, the governance framework is not complete enough for production.

👉 Read our full editorial: Evaluation-driven AI agents need specifications, evals and guardrails



   
ReplyQuote
Share: