Join our Newsletter — 33% off our NHI Course
Home› FAQ› AI Security› What is the difference between pass@k and sec_pass@k…
AI Security

What is the difference between pass@k and sec_pass@k in secure code evaluation?

← Back to all FAQ
By NHI Mgmt Group Editorial Team Updated September 30, 2026 Domain: AI Security

pass@k measures how many generated solutions satisfy functional tests. sec_pass@k measures how many satisfy both functional tests and security tests. The difference matters because code can be correct yet still vulnerable. For AI coding systems, teams should look at both metrics together to understand whether a model is producing usable code that also resists realistic attacks.

How pass@k and sec_pass@k measure different kinds of “good” code

pass@k tells you whether at least one of the generated candidates works from a functional standpoint. It is a utility metric: does the code satisfy the intended task, compile, and pass the visible tests? sec_pass@k adds a second filter, asking whether the same candidate also survives security-focused evaluation, so the result is not just correct but also materially safer.

The practical difference is that pass@k can reward code that is functionally successful but still exposes attack surface, weak validation, unsafe deserialization, insecure defaults, or other flaws. sec_pass@k is stricter because it only counts solutions that meet both correctness and security expectations. That makes it more useful when the system will be used in real software delivery, not just benchmark completion.

For secure code evaluation, the two metrics answer different questions. pass@k estimates task completion capability. sec_pass@k estimates the overlap between task completion and security acceptability. If a model scores well on pass@k but poorly on sec_pass@k, it is producing code that looks usable in a narrow functional sense but is still risky to deploy.

Why the gap between the two metrics matters

The gap between pass@k and sec_pass@k is often the most informative part of the evaluation. A large gap suggests the model can solve the programming problem but does not reliably internalize security constraints. In practice, that means security review will still need to catch the kinds of defects that the functional benchmark ignores.

A smaller gap is more encouraging because it indicates the model is not trading security away for functional success. Teams should interpret that as stronger evidence that the model can produce code that is both useful and defensible. For AI coding systems, this matters because downstream users often assume a “passing” solution is also safe enough to trust.

When comparing models, pass@k alone can overstate readiness for production assistance. sec_pass@k helps separate code-generation capability from secure-code capability, which is a better fit for environments where correctness and attack resistance both matter. The measure is especially helpful when the model is used to generate snippets, scaffolding, or patches that may be copied directly into real systems.

How to read the metrics in practice

pass@k is best treated as a baseline capability measure, not a security endorsement. sec_pass@k is a stricter operational signal, but it still depends on what security tests were used and how well they represent realistic abuse cases. A model can look strong if the security checks are shallow, or look weak if the checks are unusually narrow or overly strict.

That is why secure code evaluation should define both the functional test set and the security test set with care. The right comparison is not “does the model ever pass?” but “does it pass while avoiding the classes of weaknesses we actually care about?” That distinction is what makes sec_pass@k more decision-relevant for secure development workflows.

For this reason, teams should use the two metrics together rather than choosing one. pass@k shows whether the system can solve the task. sec_pass@k shows whether it can do so without introducing obvious security debt. The pair gives a more complete view of whether an AI coding tool is fit for constrained use, code review assistance, or guarded automation.

Practitioner Guidance

What to prioritise: Treat pass@k as a capability indicator and sec_pass@k as the deployment-relevance indicator. If the scores diverge materially, assume the model still needs security guardrails, review, or restricted use even when it appears functionally strong.

What to verify: Check that the security test suite covers the vulnerabilities you would actually block in production, not just generic style or lint issues. A useful sec_pass@k result should reflect realistic exploitability, not an artificial pass condition.

Practitioner takeaway: The main lesson is that functional correctness is necessary but not sufficient, and sec_pass@k is the metric that tells you whether the model’s “working” output is also safe enough to trust.

Deepen Your Knowledge

Sign up to our weekly newsletter — get 33% off our NHI Foundation Level Course

    NHIMG Editorial Note
    Reviewed and updated by the NHIMG editorial team on September 30, 2026.
    NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org