TL;DR: California’s new AI laws take effect on January 1, 2026 and require companion and healthcare-focused systems to prevent self-harm content, avoid misleading medical authority claims, and intervene in live conversations, according to Lakera. The shift is from policy intent to runtime control, where governance must hold up under user interaction, not just documentation.
At a glance
What this is: This analysis explains how California’s January 2026 AI laws move governance from policy documents to runtime controls for user-facing systems.
Why it matters: IAM, NHI, and AI governance teams need to treat live conversational behaviour as a controllable production surface, not just a policy or model-training issue.
Context
California’s new AI rules focus on user-facing systems that converse with people in real time, not on how models were trained or how much documentation exists. The practical problem is runtime behaviour: what the system says, when it intervenes, and whether it can hold boundaries under live interaction.
For identity and access teams, that is a governance shift as much as a safety one. AI systems that answer questions, sustain long conversations, or imitate expert guidance now need controls that operate at the moment of response, which makes runtime policy enforcement a production requirement rather than a back-office review.
Key questions
Q: What breaks when AI agents are not governed at runtime?
A: Without runtime governance, an agent can shift behaviour after provisioning and still execute actions that were never reviewed in context. That is where tool chaining, MCP connections, and rapid decision-making become dangerous. Static approval cannot stop a live change in intent, so teams lose control at the point of action.
Q: Why do user-facing AI systems need runtime guardrails for California compliance?
A: Because California is regulating observable behaviour in live conversations, not the model’s training history. If a system can imply authority, miss a crisis signal, or fail to disclose its nature during interaction, compliance fails where the user experiences the product.
Q: How should teams decide when a chatbot needs intervention logic?
A: Use intervention logic whenever a conversation can drift into self-harm, dependency, or other high-risk guidance. The trigger should be based on the content and context of the exchange, not on whether the system was originally designed as a support bot.
Q: What is the difference between disclosure and behavioural control in AI governance?
A: Disclosure tells the user what the system is, while behavioural control limits what the system can do in a live session. Both matter, but disclosure alone does not stop harmful output, misleading medical framing, or unsafe escalation once the conversation is under way.
Technical breakdown
Runtime guardrails for conversational AI
Runtime guardrails are controls that inspect, approve, block, or rewrite model output after inference and before the user sees it. They differ from training-time alignment because they can enforce context-specific policy in the moment, including crisis-response logic, disclosure language, and domain restrictions. In California’s model, the system must behave safely during the conversation itself, which means policy has to be enforced on live prompts and outputs rather than assumed from model tuning or static documentation. This is especially relevant when the system sustains long-running interactions, where risk can appear only after multiple turns and subtle escalation.
Practical implication: place policy enforcement between model output and user delivery, not only in model development workflows.
Companion chatbot disclosure and intervention logic
Companion systems create a behavioural risk because users may attribute human intent, empathy, or authority to software that is designed to stay engaged. The regulatory expectation is not just to label the system once, but to keep disclosure active and to intervene when conversation content shifts into self-harm or other high-risk territory. That implies a ruleset that tracks context over time, recognises sensitive intent, and changes the response pattern immediately. The control problem is not whether the model can generate a safe sentence, but whether the system can reliably shift from open-ended conversation to constrained intervention when the dialogue crosses a threshold.
Practical implication: define crisis triggers and intervention flows that override normal conversational behaviour once risk signals appear.
Medical authority claims in health-adjacent AI
Health and wellness systems create a different failure mode: they can sound authoritative without being licensed, factual, or clinically supervised. California’s rules target that gap by focusing on titles, phrasing, and design cues that imply medical expertise. In practice, this means governance must cover linguistic patterns and presentation choices, not just factual accuracy. A system can be technically correct and still violate policy if it presents itself as a doctor, clinician, or similar authority. The technical lesson is that trust signals are part of the attack surface, because users react to voice and framing before they evaluate source provenance.
Practical implication: review response wording and interface cues as compliance controls, not merely UX decisions.
NHI Mgmt Group analysis
Runtime control is now the compliance boundary for user-facing AI: California’s approach treats live response behaviour as the point where governance succeeds or fails. Policy decks and model documentation do not satisfy that test if the system still produces unsafe or misleading content during an actual conversation. For practitioners, this shifts control ownership from design-time intent to production-time enforcement.
Disclosure alone is not enough when users stay engaged: Companion systems are risky because repeated interaction can create perceived trust, not just awareness. California’s rules recognise that one-time notices decay over time, so disclosure has to persist while the conversation continues. The implication is that governance must account for user state, not just system state.
Medical authority is a presentation problem as much as a content problem: The AB 489 logic shows that language, tone, and interface cues can create false expertise even without explicit claims. That widens the governance surface beyond factual correctness. Practitioners need to treat trust signals as regulated behaviour, not cosmetic detail.
Runtime guardrails expose the gap between intent and enforcement: This article’s underlying concept is a runtime governance gap, where good policy exists but the live system still decides what leaves the model boundary. That gap matters because regulators now care about operational behaviour, not aspirational controls. Teams should expect more rules written against observable output, not internal design claims.
What this signals
Runtime governance gap: California’s rules show that AI programmes fail when they stop at policy language and never enforce the policy at the point of response. For user-facing systems, the control surface is the conversation itself, so teams need production guardrails that can intercept unsafe behaviour before users act on it.
For identity and access leaders, this is also a lifecycle issue for AI systems: deployment is no longer the end state, and post-release behaviour must be monitored as continuously as access. The practical signal is simple, if a system can speak to users, it must also be governable while it speaks.
For practitioners
- Implement runtime output controls Place policy enforcement after model generation and before user delivery so unsafe, misleading, or out-of-scope responses can be blocked in real time.
- Define crisis intervention triggers Map self-harm and other high-risk conversational signals to deterministic intervention flows that override normal chatbot behaviour without waiting for manual review.
- Audit trust-signalling language Review prompts, labels, interface copy, and response templates for phrases or design cues that imply clinical or human expertise.
- Instrument compliance reporting Log when guardrails trigger, which policy branch fired, and how often safety interventions occur so reporting is based on observed runtime behaviour.
Key takeaways
- California’s AI laws are pushing governance from documentation into live response controls for companion and healthcare-oriented systems.
- The core risk is not model capability alone but unsafe influence at the moment a user is already engaged with the system.
- Teams that can enforce runtime boundaries, disclosure, and intervention logic will be better positioned for this regulatory shift.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Agentic AI Top 10 addresses the attack surface, NIST AI RMF sets the technical controls, and ISO/IEC 42001:2023 and EU AI Act define the regulatory obligations.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP Agentic AI Top 10 | ASI09 — Human-Agent Trust Exploitation | The article focuses on user trust and misrecognised authority in live AI conversations. |
| ASI03 — Identity & Privilege Abuse | User-facing AI is being governed through what it is allowed to say and do at runtime. | |
| Recommendation — Constrain trust-signalling outputs so users do not mistake the system for a human or licensed expert. Enforce output boundaries that prevent the system from claiming privileges or authority it does not have. | ||
| NIST AI RMF | MANAGE — AI risk management | The article is fundamentally about operationalising AI governance in production. |
| Recommendation — Operationalise AI risk controls so post-deployment behaviour is monitored and corrected in production. | ||
| ISO/IEC 42001:2023 | AIMS — AI management system | California’s rules push governance from policy statements into an operating AI management system. |
| Recommendation — Embed runtime guardrails and incident reporting into the AI management system lifecycle. | ||
| EU AI Act | Art. 14 — Human oversight | The article’s control theme is live oversight of system behaviour when users interact with AI. |
| Recommendation — Design oversight controls that can intervene when AI behaviour drifts during live interaction. | ||
Key terms
- Runtime Guardrail: A control applied while an AI agent is operating, not just during configuration or review. Guardrails can block dangerous tool calls, require approval for sensitive actions, or stop data leakage before it reaches systems or users.
- Conversational AI Disclosure: Conversational AI disclosure is the practice of making it clear that a system is software, not a person, during an interaction. In regulated settings, disclosure must persist as the conversation continues, especially when users may infer empathy, authority, or clinical expertise.
- Trust Signalling: Trust signalling is the use of words, tone, labels, or visual cues that shape how a user interprets a system’s authority. In AI governance, these signals matter because they can mislead users even when the underlying model response is technically correct.
- Governance Gap: A governance gap is the distance between knowing an asset exists and being able to enforce policy on it. In identity programmes, it appears when discovery, review, and enforcement are split across different tools or teams, leaving access partially visible but not truly controlled.
Deepen your knowledge
NHI governance, agentic AI identity, and machine identity lifecycle are core topics in our NHI Foundation Level course, the industry's only accredited NHI security programme. If you are responsible for identity security strategy or NHI governance in your organisation, it is worth exploring.
Published by the NHIMG editorial team on June 10, 2026.
Updated on October 10, 2026.
NHI Mgmt Group, the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org