Subscribe to the Non-Human & AI Identity Journal
Home Glossary AI Security Model-level risk
AI Security

Model-level risk

← Back to Glossary
By NHI Mgmt Group Updated August 15, 2026 Domain: AI Security

Model-level risk is the unsafe behaviour embedded in an AI model itself, regardless of where it is hosted. It includes jailbreak susceptibility, misalignment, and policy behaviours that persist across deployments and must be tested separately from infrastructure controls.

Expanded Definition

Model-level risk refers to unsafe or unreliable behaviour that originates inside the AI model’s learned parameters and instruction-following patterns, not in the cloud, application, or endpoint where it runs. That distinction matters because the same model can display the same failure mode across multiple deployments, even when the surrounding infrastructure is well secured. In practice, model-level risk covers jailbreak susceptibility, harmful instruction adherence, policy evasion, hallucination patterns that create operationally dangerous outputs, and misalignment between model responses and the organisation’s intended guardrails.

For security and governance teams, the key question is whether the model itself can be induced to behave in ways that defeat safety expectations. This is why model evaluation must go beyond access control, logging, or network segmentation and include red teaming, behaviour testing, and controlled prompt abuse scenarios. The NIST Cybersecurity Framework 2.0 is useful here as a governance anchor, but it does not replace model-specific safety assessment. Definitions vary across vendors when they describe "model risk," so NHIMG treats model-level risk as a distinct class of AI security concern rather than a broad synonym for AI risk overall.

The most common misapplication is treating model-level risk as an infrastructure problem, which occurs when teams assume patching the host environment will fix behaviours that are embedded in the model itself.

Examples and Use Cases

Implementing model-level risk controls rigorously often introduces workflow friction, requiring organisations to balance stronger safety validation against slower release cycles and more limited model flexibility.

  • A customer support assistant can be prompted into revealing restricted internal guidance, showing that the model’s instruction hierarchy is weaker than expected even when deployed behind strong authentication.
  • A code-generation model may repeatedly produce insecure patterns unless it is evaluated with abuse-case prompts and domain-specific safety tests before release.
  • A procurement chatbot may comply with deceptive user requests in a way that bypasses policy intent, demonstrating persistent misalignment across environments.
  • A regulated enterprise may compare multiple model versions and find that one retains jailbreak susceptibility after fine-tuning, which means the issue is in the model behaviour rather than the application wrapper.

Teams often document these behaviours using repeatable test suites, curated prompt sets, and red-team findings aligned to NIST Cybersecurity Framework 2.0 governance practices, while also applying model-specific evaluation methods to capture unsafe outputs that standard infrastructure scans will never see. The important use case is not only identifying whether the model can fail, but whether the same failure can be reproduced after redeployment, version changes, or vendor migration.

Why It Matters for Security Teams

Security teams need to understand model-level risk because it changes how AI systems are governed, tested, and approved. If a model can be manipulated into unsafe or noncompliant behaviour, perimeter controls, IAM, and host hardening may all be correctly implemented while the actual business workflow still produces dangerous outcomes. That creates a blind spot for incident response, vendor due diligence, and approval gates. For NHI and agentic AI environments, the stakes are even higher because an autonomous agent may execute tool calls, write data, or trigger downstream workflows based on a model output that was never trustworthy to begin with.

This is why model-level risk belongs in procurement reviews, pre-production assurance, and ongoing monitoring, not only in security architecture diagrams. It also helps explain why model governance must be separated from platform governance, even when the same vendor hosts both. In practice, security teams should treat the model as an object of assessment in its own right, with evidence collected for misuse resistance, refusal behaviour, and response consistency under stress. Organisations typically encounter the consequences only after a jailbreak, unsafe automation, or policy breach, at which point model-level risk becomes operationally unavoidable to address.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

OWASP Agentic AI Top 10 and CSA MAESTRO address the attack and risk surface, while NIST AI RMF, NIST AI 600-1 and NIST CSF 2.0 set the governance and control requirements practitioners need to meet.

FrameworkControl / ReferenceRelevance
NIST AI RMFAI RMF frames governance for AI risks arising from model behaviour and misuse.
NIST AI 600-1The GenAI profile addresses risks from model outputs, misuse, and unsafe behaviour.
OWASP Agentic AI Top 10Agentic AI guidance highlights prompt abuse and unsafe model behaviour in execution contexts.
CSA MAESTROMAESTRO covers security concerns for agentic systems where model behaviour drives actions.
NIST CSF 2.0GV.RM-01CSF governance risk management supports formal treatment of AI-related operational risk.

Use AI RMF to assign ownership, assess model behaviour, and document risk treatments before deployment.

NHIMG Editorial Note
Reviewed and updated by the NHIMG editorial team on August 15, 2026.
NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org