Join our Newsletter — 33% off our NHI Course
Home Glossary AI Security Systemic-Risk Model
AI Security

Systemic-Risk Model

← Back to Glossary
By NHI Mgmt Group Updated September 6, 2026 Domain: AI Security

A systemic-risk model is a GPAI model that crosses the EU AI Act's higher-risk threshold, generally associated with training compute above about 10²⁵ FLOPs. These models face additional obligations beyond baseline documentation, including structured red-teaming, incident reporting, risk management, and cybersecurity controls for model weights and infrastructure.

Expanded Definition

A systemic-risk model is a general-purpose AI model that crosses the EU AI Act’s higher-risk compute threshold, which is generally associated with training runs above about 10²⁵ FLOPs. The term is less about everyday model capability and more about whether a model’s scale makes its failures, misuse, or infrastructure dependencies significant enough to trigger heightened governance.

In practice, the label matters because the regulatory status attaches to a model class, not just to a deployment. That means the organisation must think about lifecycle controls around the model itself: documentation, red-teaming, incident handling, and protection of model weights and supporting systems. Definitions and implementation details continue to evolve, so practitioners should read the term as a regulatory risk category rather than a product description.

For the underlying governance frame, the EU AI Act is the primary authority on when obligations escalate. The exact threshold and evidence expectations may depend on how training compute is measured and documented, which is why boundary-setting is often a reporting issue as much as a technical one.

Examples and Use Cases

Systemic-risk models appear in environments where foundation models are trained at very large scale and then reused across multiple products or business lines. The same model may support customer-facing assistants, internal automation, and developer tooling, which increases the impact of one control failure.

  • A provider trains a frontier model once and then exposes it through multiple APIs, making weight security and incident response shared concerns across all consuming teams.
  • A company fine-tunes a very large model for enterprise workflows and must decide whether the base model already meets the systemic-risk threshold before release.
  • A model lab runs structured red-teaming to test unsafe outputs, jailbreak susceptibility, and harmful tool use before deployment.
  • An organisation prepares incident reporting and governance records so it can explain model behavior, retraining events, and significant failures to regulators or customers.
  • A security team protects model weights, training infrastructure, and privileged access paths because compromise of any one layer can affect all downstream deployments.

The tradeoff is that stronger assurance usually increases coordination overhead. More review, more testing, and more evidence collection can slow release, but that cost is part of the control expectation for models whose failure would be difficult to contain.

Security Implications

Once a model falls into this category, poor security is no longer a local engineering issue. Weak protection of training assets, incomplete logging, or inadequate red-teaming can create exposure that spreads across many products and users because the same model often sits behind multiple services.

The most important failure modes are compromise of model weights, tampering with training or evaluation pipelines, and gaps in incident detection. If an attacker, insider, or third-party dependency can alter a high-value model or its surrounding infrastructure, the result may be silent integrity loss rather than an obvious outage. That is especially dangerous because model failures can propagate through automated decisioning, agentic workflows, or customer-facing systems before anyone notices.

NHIMG research shows that 97% of NHIs carry excessive privileges, increasing unauthorised access and broadening the attack surface. For systemic-risk models, that pattern is especially relevant where service accounts, API keys, and deployment pipelines can reach training or inference infrastructure.

Domain and Governance Relevance

In AI governance, the term matters because it marks a point where scale changes accountability. A large model is not just a more capable model; it is a higher-consequence asset that may require formal ownership, evidence retention, and cross-functional oversight.

For NHI and machine-identity governance, the connection is concrete rather than incidental. Training clusters, model registries, artifact stores, CI/CD pipelines, inference services, and evaluation harnesses are all driven by non-human identities and secrets. If those identities are poorly scoped or rotated, the model lifecycle itself becomes harder to trust, because the systems that create and update the model are also part of the risk surface.

That is why systemic-risk models should be read as both an AI governance issue and an identity governance issue. The model may be the regulated object, but the control failure often starts in the machine access paths that surround it.

Risk and Threat Considerations

Systemic-risk models create concentrated exposure: one compromise, one poisoned training path, or one weak privileged account can affect many downstream services at once. The threat is not only malicious misuse of the model output, but also integrity failure in the model supply chain and the infrastructure that supports it.

Failure mechanism: Adversaries or insiders can abuse over-privileged machine identities, weakly controlled training environments, or insufficiently monitored weight repositories to tamper with model assets, exfiltrate sensitive artifacts, or persist across retraining and deployment cycles.

Impact: The organisation can lose trust in the model’s outputs, expose sensitive weights or data, disrupt regulated AI operations, and propagate a single control failure across multiple products or business units.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

MITRE ATT&CK address the attack surface, NIST AI RMF and CIS Controls v8 set the technical controls, and EU AI Act and ISO/IEC 42001:2023 define the regulatory obligations.

FrameworkControl / ReferenceRelevance
EU AI ActArticle 51 — Systemic-risk GPAI modelsDefines the systemic-risk model threshold and added obligations for high-impact GPAI.
Recommendation — Classify qualifying models under Article 51 and apply the required heightened governance and reporting controls.
ISO/IEC 42001:20234.1 — Understanding the organization and its contextFrames AI governance context and accountability for high-risk model oversight.
Recommendation — Set AI governance scope so systemic-risk models have assigned ownership and documented oversight.
NIST AI RMFGOV — GovernCovers AI risk governance, documentation, and accountability for model lifecycle decisions.
Recommendation — Establish governance for model risk acceptance, documentation, and escalation.
CIS Controls v85 — Account ManagementApplies to service accounts and privileged access used in model training and deployment.
8 — Audit Log ManagementSupports detection and evidence retention for model operations and incident response.
Recommendation — Restrict and review non-human accounts that can reach model weights or training infrastructure. Log model training, access, and release activity so assurance evidence is available for review.
MITRE ATT&CKT1552 — Unsecured CredentialsMaps to secrets and keys that protect model infrastructure and artifacts.
Recommendation — Hunt for exposed credentials that could unlock model systems or artifact stores.

Practitioner Guidance

Governance implication: Treat systemic-risk classification as a boundary-setting decision, not a labelling exercise. The classification should determine who owns model assurance, which evidence must exist, and when the model is subject to heightened review before release or material change.

What to watch for: Pay close attention to shared training and deployment identities, opaque compute accounting, and missing records for red-teaming or incident handling. Those are common signs that the organisation is relying on the model’s scale without having the controls that scale demands.

Practitioner takeaway: If the model is large enough to trigger the systemic-risk threshold, the surrounding identity and infrastructure controls need to be managed as part of the model’s security posture, not as separate operational plumbing.

Deepen Your Knowledge

Sign up to our weekly newsletter — get 33% off our NHI Foundation Level Course

    NHIMG Editorial Note
    Reviewed and updated by the NHIMG editorial team on September 6, 2026.
    NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org