TL;DR: The NAIC model bulletin on AI is guidance until states adopt it, but 25 states had formally adopted it by July 2026 and another 8 were in progress, while examiners still expect documented testing, monitoring, and vendor accountability, according to Openlayer. For insurers, the compliance burden is shifting from policy language to auditable evidence that AI systems were tested before deployment and monitored after.
At a glance
What this is: This is an analysis of the NAIC model bulletin on AI and the evidentiary controls insurers need before and after deployment.
Why it matters: It matters because insurers need governance structures that can prove accountability, fairness, and monitoring across AI, third-party models, and the identity of the people or committees responsible for acting on findings.
By the numbers:
- As of July 2026, 25 states have formally adopted the NAIC Model Bulletin, with another 8 actively moving through legislative or regulatory approval processes.
👉 Read Openlayer's analysis of the NAIC AI model bulletin and insurer obligations
Context
The primary issue here is not whether insurers use AI, but whether they can prove that AI-driven decisions were tested, monitored, and governed with clear authority. In practice, the NAIC model bulletin turns AI oversight into an evidence problem, especially when carriers rely on third-party models or operate across multiple state regimes with different adoption timelines.
That governance problem has a strong identity angle because accountability depends on named owners, review committees, and the ability to trace decisions back to responsible people. For IAM, PAM, and NHI programmes, this is the same pattern seen in machine and service identities: if control ownership is vague, the organisation cannot answer for what the system did or who was empowered to stop it. The bulletin therefore behaves less like a policy statement and more like an operational control test.
The state-by-state adoption pattern is typical of US insurance regulation. The operating challenge is not a single compliance cliff, but uneven expectations, market conduct exams, and documentation gaps that can surface before formal adoption occurs.
Key questions
Q: Where do AI governance programmes fail when regulators expect examinable evidence?
A: They fail when teams rely on policy documents without versioned test results, decision logs, and named ownership. Regulators need to see who approved the model, what fairness checks were run, what threshold was used, and how failures were handled. If those artefacts do not exist, governance cannot be demonstrated during examination.
Q: Why do third-party AI models still create compliance obligations?
A: Third-party AI does not remove deployer responsibility. If an organisation uses external models, copilots, or agent workflows inside its own products or operations, it still needs oversight, logging, risk management, and evidence of accountability. The obligation follows the use case and the workflow, not just who built the model.
Q: How should insurers test whether AI decisions are unfairly discriminatory?
A: Run subgroup-level adverse impact analysis across protected classes, measure selection rates, and compare them against a defined threshold such as the four-fifths rule. Then validate whether proxy variables or calibration gaps are driving the disparity. Testing must be written up in a form regulators can inspect, not just stored in analyst notes.
Q: Who is accountable when an AI model affects a consumer decision under the bulletin?
A: The insurer is accountable, even when the model came from a third party. Regulators expect a named owner or committee with authority to govern the system, monitor outcomes, and act on adverse findings. If authority is unclear, the organisation cannot show effective oversight of the decision process.
Technical breakdown
How the NAIC AI bulletin turns guidance into examinable evidence
The NAIC model bulletin starts as guidance, but it becomes operationally binding when a state adopts it through its own regulatory process. That matters because regulators are not looking for a policy summary, they are looking for artefacts that show AI systems were selected, tested, monitored, and retired under a governed process. The AIS program requirement creates an evidence chain, not just a governance intent. For insurers, the technical challenge is maintaining versioned records, named ownership, and decision logs that can survive examination across multiple states with different interpretations.
Practical implication: build the AI inventory, review trail, and approval evidence as operational controls, not as compliance paperwork.
Bias testing, demographic parity, and the four-fifths rule
The bulletin’s fairness expectations rely on measurable outcome testing, not broad assertions that a model is unbiased. In practice, that means measuring whether protected classes experience materially different selection rates, then flagging cases where the gap exceeds the chosen threshold. The EEOC four-fifths rule is a common reference point for this kind of review because it gives examiners and insurers a concrete test for disparate impact. The key point is that fairness is evaluated through outcomes across subgroups, not through model accuracy averaged across the full population.
Practical implication: require subgroup-level fairness reports before promotion and after material model changes.
Third-party model governance and the look-through problem
The bulletin does not let insurers outsource accountability to the vendor that built the model. That is the look-through problem: regulators assess the insurer’s decision process, even when the scoring engine or claims model came from a third party. Practically, this means the insurer needs enough documentation to understand intended use, training data sources, limitations, and performance across subgroups. Without that evidence, vendor procurement becomes a governance gap because the insurer cannot demonstrate how model outputs were controlled in production.
Practical implication: align procurement, contract terms, and runtime monitoring so third-party models remain examinable after deployment.
Threat narrative
Attacker objective: The objective is not direct compromise but regulatory failure, unfair decision exposure, and ungovernable AI use that cannot be defended under examination.
- Entry occurs through insurer reliance on a third-party AI model or internal decision system without sufficiently examinable governance artefacts.
- Escalation happens when the organisation cannot show who owned the model, what testing was performed, or how monitoring results were acted on.
- Impact is regulatory exposure, adverse decision risk, and weak defensibility during state examination.
NHI Mgmt Group analysis
AI governance in insurance is becoming an evidence discipline, not a policy discipline. The bulletin does not reward generic statements of intent; it rewards auditable proof that systems were tested, monitored, and controlled. That changes the compliance conversation from who wrote the policy to who can produce the artefacts. For practitioners, the governance model now has to behave like a control system with records, thresholds, and accountable owners.
Third-party AI does not reduce insurer accountability, it relocates the burden of proof. The look-through problem means insurers remain responsible for decisions made by vendor-supplied models, even when internals are hidden. That is especially relevant where AI outputs influence underwriting, pricing, or claims decisions that touch regulated consumer rights. The practical conclusion is that procurement, runtime monitoring, and examination readiness must be designed together, not sequenced as separate workstreams.
Identity and authority are now part of AI compliance, not just access management. The bulletin assumes someone with real authority can act on adverse findings, which makes named ownership and approval scope operationally critical. In IAM terms, this resembles privileged governance over a high-impact system: if the reviewer cannot act, the control is hollow. The result is that governance roles, review committees, and escalation paths become examinable identity controls.
Regulatory fragmentation is the new AI risk multiplier for insurers. A single control design will not map cleanly across states that adopted the bulletin verbatim, modified it, or are enforcing it informally through market conduct exams. That raises the value of control mapping, evidence normalisation, and common testing thresholds. For practitioners, the market signal is clear: build once for the strictest likely examination standard, then localise only where law requires it.
Model monitoring is becoming the operational layer that policy teams cannot fake. Documentation alone can describe governance, but it cannot prove a model stayed within fairness thresholds after deployment. That is where runtime monitoring, versioned audit trails, and promotion gating matter most. The broader lesson is that AI governance is converging with security operations: if you cannot show the system behaved as expected at the time of decision, you do not have control.
What this signals
AI governance will increasingly be judged by operational proof, not by the sophistication of the policy set. For insurer programmes, that means the control question shifts to whether monitoring, promotion gating, and reviewer authority are embedded in the workflow. The most useful alignment is to treat AI governance like a regulated identity system, where ownership, decision rights, and evidence trails matter as much as the model itself.
Control fragmentation across states creates a versioning problem for enterprise governance. The practical response is to standardise on one internal evidence model, then map it to the strictest state expectation rather than maintaining separate control narratives for each jurisdiction. For AI teams, this is similar to maintaining a single source of truth for identity entitlements while handling local policy exceptions.
Identity governance and AI governance are converging around the same failure mode: authority without proof. A named approver who cannot demonstrate actionability is no better than an unmonitored service account with standing privilege. For readers building cross-domain programmes, the implication is to integrate IAM, model risk, and audit evidence so each decision can be traced to a responsible actor and a verifiable control.
For practitioners
- Create a state-mapped AI inventory Maintain a live inventory of every underwriting, claims, pricing, and fraud model, then map each system to the states where Bulletin expectations apply or are being enforced informally. Include model owner, business purpose, data sources, approval date, and last review date.
- Require versioned fairness evidence before promotion Block deployment until each model has subgroup-level testing records, threshold definitions, reviewer sign-off, and a version hash tied to the exact artefact promoted. Preserve those records so examiners can inspect the decision trail after the fact.
- Close the look-through gap in vendor contracts Add audit rights, incident notification duties, evidence-sharing obligations, and model documentation requirements to vendor agreements. If the vendor cannot support examination requests for training data, limitations, and subgroup performance, treat that as a control failure rather than a procurement issue.
- Assign authority to a named AIS owner Designate a person or committee with explicit authority to stop, escalate, or retrain models when monitoring reveals adverse impact or drift. Document the escalation path so the role is not symbolic and can actually act on findings.
Key takeaways
- The NAIC bulletin turns AI oversight into an examinable control problem, not a policy exercise.
- Insurers remain accountable for third-party model outcomes, so vendor reliance does not remove the need for runtime evidence.
- Programmes that combine named ownership, fairness testing, and versioned audit trails will be better positioned for multi-state examination.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
NIST AI RMF and NIST CSF 2.0 set the technical controls, while GDPR define the regulatory obligations.
| Framework | Control / Reference | Relevance |
|---|---|---|
| NIST AI RMF | GOVERN | The bulletin centres accountability, governance, and oversight for AI systems. |
| GDPR | Art.22 | AI-driven insurance decisions can affect individuals through automated decision-making. |
| NIST CSF 2.0 | PR.DS-1 | Model evidence and data lineage need protection as governed assets. |
Use AI RMF GOVERN to assign ownership, document review rights, and keep AI governance examinable.
Key terms
- Artificial intelligence system program: A formal governance structure for selecting, validating, monitoring, and retiring AI systems. In the insurance context, it has to produce evidence that the model was controlled before deployment and after changes, not just documented in policy language. The program becomes examinable only when records, owners, and thresholds are operationally complete.
- Adverse impact analysis: A fairness test that checks whether a protected group experiences worse outcomes than a reference group. In practice, insurers use it to evaluate selection rates, approval rates, or adverse action rates across subgroups, then compare the results against a defined threshold that can be defended during review.
- Look-through problem: The governance failure that occurs when an organisation assumes a vendor relationship transfers accountability away from the buyer. In regulated AI use, the insurer still owns the outcome and must show how third-party models were tested, monitored, and controlled in production. The contract does not replace the control.
- Versioned audit trail: A record set that links a specific model version to its tests, approvals, thresholds, and reviewer actions. This matters because regulators and auditors inspect what was true at the time of decision, not what the programme later claims it intended to do. Without versioning, evidence loses probative value.
What's in the full article
Openlayer's full article covers the operational detail this post intentionally leaves for the source:
- A state-by-state adoption breakdown of the NAIC model bulletin and how those differences affect exam readiness.
- Examples of the exact AIS programme artefacts examiners expect, including inventory records, fairness tests, and reviewer logs.
- A deeper explanation of the AI Systems Evaluation Tool and how it changes what regulators inspect in practice.
- Specific guidance on model promotion gating, audit trail generation, and evidence retention for insurer workflows.
Deepen your knowledge
The NHI Foundation Level course, the industry's only accredited NHI security programme, covers NHI governance, workload identity, and secrets management for practitioners who need to connect control design to operational accountability. It is relevant for teams building identity governance across human, machine, and AI-adjacent programmes.
Published by the NHIMG editorial team on August 2, 2026.
NHI Mgmt Group — the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org