Sec Pass@k measures the percentage of generated solutions that pass both correctness tests and security tests within k attempts. It is a stricter benchmark than pass@k because it requires code to work properly and resist the exploit checks applied to the task.
What Sec Pass@k Measures in Practice
Sec Pass@k is a benchmark for evaluating generated code under a dual requirement: it must solve the task correctly and also survive explicit security checks. The metric is useful when a model can produce working code that is still unsafe, brittle, or exploitable.
That makes the term broader than “does it compile” or “does it pass unit tests.” Sec Pass@k asks whether security quality is present in the same candidate solution, not as a separate afterthought.
How Sec Pass@k Differs from pass@k
Traditional pass@k measures the chance that at least one of k generated attempts passes functional evaluation. Sec Pass@k tightens that bar by requiring the candidate to clear both correctness and security gates within the same attempt window.
This difference matters because a model may learn to optimize for surface correctness while still emitting unsafe patterns such as broken access checks, injection-prone handling, unsafe defaults, or secret exposure. Sec Pass@k is therefore a stricter measure of practical usefulness for security-sensitive code generation.
What Sec Pass@k Reveals About Model Quality
Sec Pass@k is best read as a joint indicator of solution quality and security robustness. A strong score suggests the model can repeatedly generate code that is not only functionally acceptable, but also resistant to the exploit checks used in the evaluation.
Because the metric combines two gates, it highlights a common gap in software generation systems: correctness can improve faster than security hardening. That makes Sec Pass@k especially relevant for code assistants, secure-by-default generation, and any workflow where generated output may enter production with minimal review.
In practice, the metric is only as meaningful as the underlying tests. If the functional checks are weak or the exploit checks are narrow, Sec Pass@k can overstate robustness. If the tests are well designed, it becomes a more realistic proxy for whether generated code can satisfy both implementation and defensive expectations.
Where Sec Pass@k Is Most Useful
Sec Pass@k is most useful when teams need a single benchmark that compares generations on both utility and safety. It helps separate models that can write plausible code from models that can produce code fit for security-aware use cases.
It is also a helpful research metric when studying trade-offs between generation diversity, repair strategies, and security alignment. For a practical explanation of how this kind of benchmark fits into broader security and control thinking, see NIST Cybersecurity Framework 2.0, OWASP API Security Top 10, and SLSA.
Risk and Threat Considerations
Sec Pass@k can create false confidence if teams treat a benchmark score as proof of real-world safety. A model may pass curated exploit tests while still producing code that is vulnerable under different inputs, environments, dependency states, or abuse paths.
Failure mechanism: The evaluation only measures the security conditions encoded in the benchmark, so narrow test coverage, weak adversarial cases, or domain mismatch can let insecure code look acceptable. That gap is especially dangerous when the generated code will handle authentication, authorization, data handling, or externally reachable interfaces.
Impact: Teams may deploy code that is functionally correct but still exploitable, leading to injection risk, privilege abuse, data exposure, or unsafe API behavior after release.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP ASVS, NIST SP 800-53 Rev 5, CIS Controls v8, OWASP SAMM and SLSA set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP ASVS | V4 — API and Web Service | Sec Pass@k often evaluates whether generated code survives security checks on service and API behavior. |
| Recommendation — Verify generated service code against API and web-service security requirements before promotion. | ||
| NIST SP 800-53 Rev 5 | SI-10 — Information Input Validation | Sec Pass@k is directly concerned with code that resists exploit inputs while remaining correct. |
| Recommendation — Apply input-validation controls to generated code that accepts external or adversarial data. | ||
| CIS Controls v8 | CIS-16 — Application Software Security | Sec Pass@k measures whether software outputs remain secure enough for real application use. |
| Recommendation — Use application security controls to assess generated code before it enters production. | ||
| OWASP SAMM | implementation — Implementation | Sec Pass@k supports secure development maturity by measuring secure output quality during code generation. |
| Recommendation — Track secure coding maturity with evaluation gates that include security as well as correctness. | ||
| SLSA | source — Build and provenance integrity | Sec Pass@k is often used alongside integrity-minded software pipelines that demand trustworthy artifacts. |
| Recommendation — Require stronger provenance and integrity checks for generated code artifacts. | ||
Practitioner Guidance
Why practitioners should care: Sec Pass@k is most valuable when it is used as a comparative signal, not a final security verdict. It should inform model selection, prompt strategy, and evaluation design, but it should not replace code review, dynamic testing, or threat modelling.
What to watch for: If a benchmark looks strong while security failures still appear in adjacent tasks, the likely problem is evaluation design rather than model robustness. The useful question is whether the exploit checks reflect the kinds of misuse the code will actually face.
Related resources from NHI Mgmt Group
Deepen Your Knowledge
Reviewed and updated by the NHIMG editorial team on September 30, 2026.
NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org