Use CVSS only when the agent behaves like a conventional application with tightly bounded actions. If the system can improvise, retain memory, or coordinate across agents, CVSS should be treated as incomplete and supplemented with agentic behaviour scoring before prioritisation.
When should a CVSS score be trusted for an AI agent?
CVSS is useful only when the AI agent behaves like a bounded application component with clear inputs, outputs, and predictable failure modes. Once the agent can plan, remember, call tools, or coordinate with other agents, the score stops capturing the full blast radius. At that point, CVSS is still a data point, but not a complete prioritisation method.
Why CVSS breaks down as agent autonomy increases
CVSS was built to describe vulnerability severity in a relatively stable technical target. An agent changes the target because the same defect may produce very different outcomes depending on tool access, memory scope, delegation chains, and whether actions are one-shot or iterative. That means a score can understate real-world risk when the agent can adapt at runtime or reuse prior context.
For tightly constrained agents, CVSS can still help compare issues such as input validation failures, authentication flaws, or exposed interfaces. For more capable systems, the more important question is not just whether a vulnerability exists, but what the agent can do with it, how far it can move, and whether one compromised action can cascade into many. That is why teams should pair severity scoring with an assessment of autonomy, authority, and control boundaries.
Teams should also treat the surrounding architecture as part of the risk story. A vulnerability that looks moderate in isolation can become severe if the agent has persistent memory, shared credentials, access to sensitive tools, or the ability to act across environments. In those cases, the score is describing the defect, not the agentic behaviour that turns the defect into impact.
What evidence makes a CVSS score more or less reliable for prioritisation?
Trust the score more when the system is single-purpose, stateless, and tightly permissioned. Trust it less when the agent can choose tools, maintain state across sessions, or influence downstream systems without fresh human review. The more the agent can improvise, the more the prioritisation problem shifts from classic vulnerability severity to runtime authority and behavioural control.
One useful test is whether the vulnerability can be exploited without any meaningful change in the agent’s decision-making. If yes, CVSS may be a fair first-pass severity signal. If exploitation depends on the agent’s ability to chain actions, retain memory, or reuse privileges, the score should be treated as incomplete and supplemented with an agent-specific review of behaviour and privilege scope.
Another practical signal is whether the agent’s outputs are reversible and easy to contain. If a bad action can be rolled back quickly, the score may align reasonably well with operational impact. If the agent can make durable changes, propagate misinformation, or trigger side effects across systems, the score will usually lag behind the real consequence.
How should teams use CVSS without letting it drive the wrong decision?
Use CVSS as a baseline technical lens, then add agentic factors that change the decision: autonomy, memory persistence, tool reach, cross-agent coordination, and the sensitivity of the actions the agent can take. This prevents teams from treating every AI-related finding as either a generic app bug or an existential incident.
When triaging, separate defect severity from agent capability. A low-scoring weakness inside a highly capable agent may deserve faster response than a higher-scoring flaw in a narrow, non-autonomous service. In practice, the right priority is the combination of exploitability, authority, and operational reach, not the CVSS number alone.
For deeper guidance on agent authority and containment, see the AI Agent Authorisation Guide, the Zero Trust for AI Agents, and the AI Agent Observability, Audit and Incident Response Guide. For agentic-risk framing beyond classic vulnerability scoring, compare findings with the OWASP Agentic AI Top 10 and NIST AI Risk Management Framework.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Agentic AI Top 10 addresses the attack and risk surface, while NIST AI RMF and NIST CSF 2.0 set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP Agentic AI Top 10 | ASI03 — Identity & Privilege Abuse | AI agent autonomy changes exploit impact through authority and privilege misuse. |
| ASI08 — Cascading Failures | Agent actions can chain across systems, making single-vuln severity misleading. | |
| Recommendation — Assess agent privilege paths and reduce standing authority before relying on CVSS priority. Model downstream cascades before deciding that a CVSS score is sufficient. | ||
| NIST AI RMF | Govern | CVSS use here depends on governance over AI risk prioritisation and escalation. |
| Recommendation — Set an AI-risk triage policy that supplements CVSS with autonomy and impact review. | ||
| NIST CSF 2.0 | GV.RM-01 — Risk Management Strategy | Teams need a risk strategy for when standard severity scoring is insufficient. |
| ID.RA-01 — Asset Vulnerabilities Are Identified and Recorded | Agentic systems need vulnerability assessment that includes their operational context. | |
| Recommendation — Define when CVSS is only a baseline and when agent behaviour scoring overrides it. Record agent capability and authority alongside each vulnerability finding. | ||
Practitioner Guidance
What to verify: Confirm whether the agent can persist state, invoke tools, or act across trust boundaries before using CVSS as a prioritisation input. If any of those are true, require an additional behavioural or privilege-based assessment.
Decision rule: If the system cannot materially change its own actions, CVSS can be a strong first-pass ranking tool. If the system can improvise or chain actions, treat CVSS as incomplete and rank by likely impact on authority, memory, and tool access.
What practitioners underestimate: The score describes a vulnerability, not the agent’s capacity to amplify it. In agentic systems, that amplification is often the real driver of urgency.
Practitioner takeaway: CVSS is trustworthy for AI agents only when the agent behaves like a tightly bounded application; as soon as autonomy expands, prioritisation must shift from raw severity to the combination of exploitability, authority, and blast radius.
Related resources from NHI Mgmt Group
- How do security teams decide whether to let AI agents automate investigations?
- How do teams decide whether an AI gateway is necessary for production agents?
- How do teams decide whether to trust an AI tool call?
- How do security teams decide whether to trust AI output in offensive or red-team workflows?