Partial-success scores matter because many AI attacks do not fail cleanly. A model may leak part of a prompt, follow only some malicious instructions, or expose enough state to make the next exploit easier. Graded scoring shows whether defenses are reducing attack impact, not just whether they are blocking the most obvious case.
Why partial-success scores change what GenAI security tests actually tell you
Partial-success scoring turns GenAI testing from a pass or fail exercise into a measurement of how much damage a weak defense still allows. In practice, many attacks are messy, incremental, or only partly effective. If you only count clean failures, you miss leakage, partial compliance, and the easier follow-on exploit paths that a determined attacker can still use.
What partial success reveals about failure modes
A graded score shows whether the model resisted the request completely, partially complied, or revealed enough to matter operationally. That distinction is important because a small amount of leaked context, a partially executed instruction, or a truncated safety failure can still expose secrets, policy logic, or hidden state. The point is not only whether the control blocked the attack, but whether it reduced the attack’s usefulness.
That matters especially in GenAI because model behavior often degrades along a spectrum rather than collapsing cleanly. A prompt injection may not produce the exact forbidden output, yet it may still steer the model, weaken refusal behavior, or surface fragments that help the next attempt succeed. Partial-success scoring captures that intermediate state instead of treating it as harmless noise.
In security testing terms, this is closer to measuring containment than measuring perfect prevention. For NIST AI 600-1 GenAI Profile, graded outcomes align with the need to assess content provenance, pre-deployment testing, and residual risk in generative systems rather than assuming binary control effectiveness.
How to use partial-success scores in a test program
Partial-success scores are most useful when they are tied to a threat model, not just a benchmark leaderboard. A test suite should define what counts as full success, partial success, and low-risk noise for the specific system being evaluated. For example, leaking one policy rule may matter less than leaking system prompt structure, but both can indicate that the defense is brittle under pressure.
A practical way to read the score is to ask three questions: did the test cause disclosure, did it change model behavior, and did it create a better launch point for the next attack? If the answer is yes to any of those, the defense is not simply “working” because it avoided the worst case. It is only partially effective, and that is exactly the signal the score is meant to preserve.
Security teams should also compare scores across versions, because a change that reduces full success but increases partial success may still be a regression. A more usable defensive model is one that leaks less, resists longer, and gives attackers less leverage. The NIST AI Risk Management Framework is helpful here because it frames AI testing as risk management, not just control verification.
Why binary pass or fail testing misses GenAI risk
Binary testing encourages false confidence. A system can appear safe if it blocks a headline attack string, even while remaining vulnerable to softer variations, multi-turn manipulation, or indirect leakage. Partial-success scoring exposes that gray zone, which is where many real compromises begin.
This is especially important for red teaming and regression testing. If a model becomes slightly harder to exploit but still reveals enough context to enable chaining, the security posture has improved only marginally. That makes the score a better indicator of resilience, because it tracks the amount of attacker progress rather than only the existence of an exploit outcome.
For broader program alignment, frameworks such as NIST Cybersecurity Framework 2.0 support this kind of outcome-based measurement by emphasizing governance, protection, detection, response, and recovery as connected functions.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
NIST AI 600-1, NIST AI RMF and NIST CSF 2.0 set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| NIST AI 600-1 | Generative Artificial Intelligence Profile | GenAI testing should measure residual risk, provenance, and pre-deployment effectiveness. |
| Recommendation — Use graded test results to track residual GenAI risk and refine pre-deployment controls. | ||
| NIST AI RMF | AI Risk Management Framework | Partial-success scoring supports risk-based evaluation of AI systems, not just binary control checks. |
| Recommendation — Assess how much attacker leverage remains after each test and treat it as risk evidence. | ||
| NIST CSF 2.0 | GV.OV — Cybersecurity Outcomes are Verified | Outcome-based testing fits verification of whether protections are actually reducing harm. |
| Recommendation — Verify that GenAI controls reduce impact, not only that they block obvious failures. | ||
Practitioner Guidance
What to prioritise: Treat partial-success scores as a ranking signal for residual exposure, not as a cosmetic metric. The most important cases are the ones where the model leaks state, weakens refusals, or creates useful attacker knowledge even if the final malicious output is not fully achieved.
What to verify: Confirm that the scoring rubric distinguishes between harmless near-misses and partial compromises that change attacker options. If a test result would help an adversary refine the next prompt, it deserves to be scored as a meaningful security outcome.
What good looks like: Scores should trend downward across releases for the same attack class, and the remaining partial successes should become less informative, less transferable, and less exploitable over time.
Practitioner takeaway: In GenAI security testing, the useful question is not only “Did we block the attack?” It is “How much attacker advantage still survived?”
Related resources from NHI Mgmt Group
Deepen Your Knowledge
Free weekly newsletter
Subscribe to the NHI & AI Identity Journal
The latest on NHI and Agentic AI security – articles, research, breaches, news and events every week.
Bonus 33% off our NHI Course when you subscribe.
Reviewed and updated by the NHIMG editorial team on October 11, 2026.
NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org