By NHI Mgmt Group Editorial TeamDomain: AI SecuritySource: PromptfooPublished December 15, 2025

TL;DR: AI regulation now reaches product teams through procurement, questionnaires, and contract evidence, with federal, state, and EU rules pushing model cards, evaluations, acceptable use policies, and incident logging into standard delivery workflows, according to Promptfoo. Documentation is no longer a governance afterthought. It has become a measurable product requirement that reshapes how AI systems are tested, sold, and operated.


At a glance

What this is: This is an analysis of how AI regulation moved into product, procurement, and evidence workflows in 2025, with model cards, evaluation artefacts, and acceptable use policies becoming standard asks.

Why it matters: It matters because security, IAM, compliance, and AI governance teams now need defensible documentation for systems that can take actions, not just generate text, and those requirements often intersect with identity, access, and audit controls.

By the numbers:

👉 Read Promptfoo's analysis of AI regulation, procurement, and compliance evidence


Context

AI regulation is no longer confined to policy teams or legal review. In 2025, procurement language, customer questionnaires, and contract evidence became the practical route by which model documentation, evaluation artefacts, and acceptable use policies reached product teams. That shift matters because AI systems increasingly behave as operational systems, not isolated models.

For practitioners, the governance gap is not the absence of rules. It is the gap between what AI systems can do and what teams can prove about those systems. Once tools, retrieval, memory, and action paths enter the stack, the same identity, access, logging, and review expectations that apply to other production systems start to apply here too. This is where the overlap with IAM, secrets management, and NHI governance becomes operational, not theoretical.


Key questions

Q: How should security teams document AI systems for procurement and compliance reviews?

A: Security teams should document the deployed AI stack, not just the base model. That means model cards, evaluation artefacts, acceptable use policies, feedback routes, and ownership details. The documentation should show prompts, retrieval sources, tools, and review points so an external buyer or auditor can understand what the system can do and how its behaviour is controlled.

Q: Why do AI agents create a governance problem for IAM teams?

A: AI agents create a governance problem because they authenticate and act as autonomous software entities with tool access. If their actions are logged only as application activity, teams lose accountability, context, and revocation clarity. IAM must therefore extend to agent identity, delegated authority, and control-plane audit trails.

Q: What breaks when AI testing ignores tools, retrieval, and memory?

A: Testing breaks down when it covers only model output and ignores the operational stack. A system can pass language tests yet still choose the wrong tool, expose sensitive retrieval data, or take unsafe actions after a context change. For governance, that means the evaluation does not match the deployed risk surface.

Q: Who is accountable when an AI system makes a harmful decision?

A: Accountability should follow the identity chain that authorized, configured, or triggered the action, including the human owner, the platform team, and any delegated agent or tool account. If the organisation cannot name that chain, the governance model is too weak for regulated AI use.


Technical breakdown

How procurement turns AI behaviour into evidence

AI procurement now treats behaviour as something that must be documented, tested, and repeated. Model cards describe what a system is, evaluation artefacts show how it behaves under test, and acceptable use policies define the boundaries of permitted use. For enterprise buyers, those artefacts become part of the control surface because they determine whether the system can be assessed consistently across vendors, business units, and deployment contexts. The important shift is that evidence is no longer optional narrative. It is the basis for contract language, risk acceptance, and auditability.

Practical implication: standardise the evidence package required before any AI system enters procurement or renewal review.

Why agentic systems create a broader control surface

The regulatory challenge grows once a system can call tools, read retrieval sources, maintain memory, or mutate external state. At that point, evaluation must cover not only output quality but also tool selection, error handling, rollback behaviour, and the identity and permissions attached to each action path. That is where AI governance intersects with IAM and NHI control, because tokens, API keys, service accounts, and delegated access define what the system can actually do. A text model may be harmless in isolation, but an agent with standing privileges is a governance object.

Practical implication: review tool permissions and credential scope alongside model testing, not after deployment.

What documentation means when regulation follows the product stack

The article shows a policy stack that runs from executive orders to OMB memos to procurement language and then to operational evidence. That stack matters because it moves compliance from abstract policy into the build and release process. Systems that interact with untrusted input, use retrieval, or expose user-facing actions need tests that reflect the deployed configuration, not a lab-only model snapshot. For teams, that means documentation must be versioned with prompts, tools, data sources, and logging so the system can be re-tested when any of those elements change.

Practical implication: version AI documentation with the deployed stack so compliance artefacts stay aligned to production reality.


NHI Mgmt Group analysis

Documentation is becoming the control plane for AI governance. The article shows that model cards, evaluations, acceptable use policies, and feedback workflows are now procurement objects, not optional supporting material. That shifts accountability from informal assurance to auditable proof, which is exactly how governance matures in regulated environments. Practitioners should treat documentation as a first-class operational control, not a paper trail.

AI governance debt: when systems ship faster than evidence, compliance becomes retroactive. The problem is not that organisations lack policies. It is that deployed systems often outpace the artefacts needed to explain, test, and defend them. That debt grows fastest where prompts, retrieval, tools, and human review points are not tracked as part of the release process. Teams should assume every undocumented path becomes a future audit failure.

Agentic AI pulls IAM and NHI into the same conversation as model risk. Once an AI system can use tools or external services, its access model matters as much as its output quality. API keys, service accounts, and delegated tokens determine whether the system can act safely, and that is a classic identity governance problem with a new runtime. Practitioners should align AI review with identity lifecycle, secret management, and least-privilege controls.

Regulation is converging on measurable behaviour, not marketing claims. The article’s federal and state examples show a consistent pattern: regulators want testing, disclosure, and incident handling that can be evidenced by someone outside the build team. That means claims about autonomy, accuracy, or replacement value need the same discipline as security claims. Security and compliance leaders should insist that AI assurances map to repeatable tests and documented operating limits.

What this signals

AI governance debt: teams that treat documentation as a post-release task will struggle to satisfy procurement, audit, and regulatory evidence requests as AI systems become more agentic. The practical response is to version documentation with the deployed stack and keep identity, tool, and logging records in the same control set.

For identity programmes, the signal is clear: AI access now needs lifecycle ownership. When a system can act through service accounts, API keys, or delegated tokens, secret scope and revocation become part of AI governance, not just IAM hygiene. That is why teams should align AI controls with the Ultimate Guide to NHIs , Regulatory and Audit Perspectives and the NIST Cybersecurity Framework 2.0.


For practitioners

  • Build an AI evidence pack for procurement Create a standard package that includes model cards, evaluation results, acceptable use policies, incident handling steps, and owner contacts before the system is offered to buyers or internal customers.
  • Test the deployed stack, not just the model Run evaluations against prompts, retrieval sources, tools, memory, and logging so the review reflects the production configuration rather than a standalone benchmark.
  • Review identity and secret scope for every AI action path Map each tool call to the API keys, service accounts, or delegated tokens it depends on, then remove standing privilege wherever the action does not require persistent access.
  • Version compliance artefacts with each release Tie documentation updates to prompt changes, tool changes, and data-source changes so the evidence remains aligned to what is actually running in production.

Key takeaways

  • AI regulation is now reaching product teams through procurement, documentation, and evidence requirements, not only formal statutes.
  • The biggest governance gap is between what AI systems can do and what teams can prove about those systems in production.
  • Identity, access, and secret controls now sit inside AI governance whenever systems can call tools or take actions.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

OWASP Agentic AI Top 10 address the attack and risk surface, while NIST AI RMF, NIST AI 600-1, NIST CSF 2.0 and NIST SP 800-53 Rev 5 set the governance and control requirements practitioners need to meet.

FrameworkControl / ReferenceRelevance
NIST AI RMFGOVERNThe article centres on AI governance, documentation, and accountability across procurement and deployment.
NIST AI 600-1The article maps directly to GenAI documentation, evaluation, and incident handling expectations.
NIST CSF 2.0GV.OV-01Governance oversight applies because the article is about evidence, policy, and operational accountability.
NIST SP 800-53 Rev 5AU-2Audit and accountability controls support the logging and evidence demands described in the article.
OWASP Agentic AI Top 10AGENT-03Agentic systems and tool use are central to the article’s discussion of deployed AI risk.

Use GenAI profile guidance to require testing artefacts, provenance, and incident procedures for deployed systems.


Key terms

  • Model Card: A structured record for one AI model that captures purpose, data sources, risk tier, ownership, approval history and known limitations. It is the primary evidence artefact that lets auditors and operators understand what a model is meant to do and who is responsible for it.
  • Verification Artefact: A verification artefact is any record created during identity proofing, including images, scores, approval notes, or vendor returns. These artefacts are valuable for audit and fraud review, but they also create privacy and breach risk if they are retained too long or exposed broadly.
  • Acceptable Use Policy: An acceptable use policy defines which data, tools, workflows, and actions are permitted for an identity or system. For AI governance, it becomes the boundary that turns vague intent into enforceable scope, which auditors and security teams can test against actual runtime behaviour.
  • Agentic AI: Autonomous AI systems capable of planning, deciding, and taking actions — including calling APIs, writing code, and orchestrating other agents — with minimal human oversight. Agentic AI introduces new NHI risks as agents must authenticate to external services.

What's in the full article

Promptfoo's full article covers the operational detail this post intentionally leaves for the source:

  • Detailed timeline of US federal, state, and EU AI requirements for 2026 planning
  • Specific procurement artefacts agencies are asking for, including model cards and evaluation evidence
  • Comparative breakdown of how federal, California, Colorado, and EU obligations differ in practice
  • Implementation implications for testing agentic systems across prompts, tools, retrieval, and logging

👉 The full Promptfoo article covers the federal timeline, state law differences, and deployment evidence requirements in more detail.

Deepen your knowledge

NHI Foundation Level course, the industry's only accredited NHI security programme, covers NHI governance, machine identity security, and secrets management. It gives practitioners a structured way to connect identity controls to the broader security and compliance programmes their organisations rely on.
NHIMG Editorial Note
Published by the NHIMG editorial team on August 1, 2026.
NHI Mgmt Group — the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org