CVSS breaks when the system can act, choose tools, and continue without human gating, because severity is no longer just a property of the code. An agent can amplify a modest flaw into a broader operational event, so teams need a behaviour-aware layer on top of traditional scoring.
Why CVSS stops being enough once an agent can act
CVSS is useful for ranking a vulnerability in isolation, but an agentic system changes the question from “how bad is the flaw?” to “what can the system do with it?” When the software can choose tools, chain actions, or continue operating without a person approving each step, a low or medium technical weakness can become a much larger operational event.
That is why CVSS alone can understate the real exposure. In an agentic setting, impact depends on autonomy, tool access, context, and the surrounding guardrails, not just the exploitability of one code path.
Behaviour-aware assessment also has to separate the defect from the workflow. A flaw that is tolerable in a passive application may be far more serious in an agent that can call APIs, modify state, retrieve data, or hand off to other tools without interruption.
What changes when the system can keep going after a mistake
In a conventional system, many scoring assumptions are relatively stable: the vulnerability is present, the attacker may exploit it, and the damage is bounded by the affected component. In an agentic system, the same weakness can cascade because the system itself can keep executing, amplify the request, or repeat the action at scale. That makes the operational blast radius a first-class part of the analysis.
This is especially important when the agent has delegated access or standing permissions across multiple tools. The question is no longer only whether the flaw exists, but whether the agent can turn that flaw into data exposure, unauthorised action, or downstream compromise before anyone notices.
Behaviour also matters for false confidence. A score may look modest if the vulnerability is local, yet the agent may use that local foothold to pursue higher-impact actions, such as retrieving secrets, invoking adjacent services, or triggering side effects in connected systems.
How to score agentic AI without losing the operational picture
Practitioners should treat CVSS as one input, not the final severity decision. Use it to describe the technical defect, then add an agent-specific layer for autonomy, tool reach, privilege, and the likelihood of repeated or chained action. That is the only way to distinguish a contained bug from a behaviour-driven incident.
For practitioners building a scoring process, the most useful question is whether the agent can convert the flaw into action. If the answer is yes, the real severity depends on what the agent can reach, whether human approval is required, and how quickly the environment can detect or stop the behaviour. Agentic AI Security Guide is a useful reference for thinking about inputs, tools, orchestration, and identity together rather than in isolation.
A practical model is to keep CVSS for code-centric severity, then add a separate control-plane review for autonomy and access. That review should answer what the agent can touch, what it can persist with, and whether one exploit can become many actions before containment. AI Agent Authorisation Guide and Zero Trust for AI Agents both support that kind of least-privilege, per-action thinking.
Risk and Threat Considerations
Agentic systems widen the impact of ordinary vulnerabilities because adversaries do not need the flaw to be severe in code terms, only useful in behaviour terms. If the agent can keep acting, the attacker may gain persistence, repeated execution, or access to downstream tools and data that the original defect never seemed to expose.
Failure mechanism: A vulnerability that would otherwise be limited to one request or one component is combined with agent autonomy, tool use, or delegated access, allowing the system to extend the attack path and create broader compromise.
Impact: The practical outcome can be unauthorized actions, data exposure, privilege abuse, or a larger operational incident than CVSS suggests, especially when the agent can repeat or chain actions before intervention.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Agentic AI Top 10 addresses the attack and risk surface, while NIST AI RMF sets the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP Agentic AI Top 10 | ASI02 — Tool Misuse | Agent tools can turn a modest flaw into harmful action. |
| ASI03 — Identity & Privilege Abuse | Severity changes when an agent can act with excess authority. | |
| ASI08 — Cascading Failures | Agentic execution can amplify one flaw into broader incident impact. | |
| Recommendation — Restrict tool use so vulnerable agents cannot chain unsafe actions. Apply per-action least privilege and human approval for privileged agent actions. Assess whether a defect can cascade across tools, sessions, or workflows. | ||
| NIST AI RMF | GOVERN — Govern | AI risk governance is needed to score behaviour, not just code severity. |
| MAP — Map | Mapping the system context reveals autonomy, tools, and dependencies affecting severity. | |
| Recommendation — Establish governance that includes autonomy and operational impact in severity decisions. Map agent capabilities, permissions, and dependencies before assigning severity. | ||
Practitioner Guidance
What to prioritise: Separate technical severity from behavioural severity. If an agent can execute the vulnerable path and then continue with tools, permissions, or memory intact, treat the issue as a control failure, not just a patching ticket.
What to verify: Check whether the vulnerable path is gated by human approval, whether tool permissions are task-scoped, and whether the agent can persist beyond the single action that triggered the issue. If those answers are unclear, the score is incomplete.
Common mistake: Teams often inherit the CVSS number from the underlying software and stop there. That misses the core question in agentic environments, which is whether the flaw can be operationalised into autonomous behaviour.
Practitioner takeaway: Use CVSS to describe the defect, but use agent behaviour to decide severity, because autonomy, reach, and repeated action are what turn a vulnerability into real exposure.
Related resources from NHI Mgmt Group
- When does just-in-time access reduce risk for agentic AI, and when does it fall short?
- How should security teams govern machine identity credentials in agentic AI environments?
- Why is identity such a critical factor in securing AI agent systems?
- When is it appropriate to implement MCP in the context of AI systems?
Deepen Your Knowledge
Free weekly newsletter
Subscribe to the NHI & AI Identity Journal
The latest on NHI and Agentic AI security – articles, research, breaches, news and events every week.
Bonus 33% off our NHI Course when you subscribe.
Reviewed and updated by the NHIMG editorial team on October 11, 2026.
NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org