Join our Newsletter — 33% off our NHI Course
Home Glossary AI Security Run-to-run variance
AI Security

Run-to-run variance

← Back to Glossary
By NHI Mgmt Group Updated September 7, 2026 Domain: AI Security

The difference in outputs produced by the same model across repeated runs on the same task. In security analysis, high variance can hide findings in one pass and reveal them in another, which is why repeated evaluation often gives a truer operational picture.

Expanded Definition

Run-to-run variance describes how much a model’s output changes when the same prompt, data, or evaluation is repeated under similar conditions. The term matters in AI security because reproducibility is not just a research concern; it affects whether a model reliably exposes harmful content, policy violations, prompt-injection effects, or false negatives across repeated tests.

In practice, low variance suggests stable behaviour, while high variance means the result can shift enough to change an analyst’s conclusion from one run to the next. That distinction is especially important when teams use a single pass to judge safety, robustness, or detection quality. A model can appear sound in one execution and fail in the next without any change to the input.

Guidance-vs-consensus note: there is no universal threshold for an acceptable level of run-to-run variance. The operational question is whether the observed spread is small enough for the use case, the control objective, and the evaluation method being used. For readers comparing evaluation discipline with broader security controls, the NIST controls catalogue provides useful context on repeatable assessment and monitoring practices: NIST SP 800-53 Rev 5 Security and Privacy Controls.

Examples and Use Cases

Run-to-run variance shows up anywhere repeated AI evaluation is expected to produce consistent evidence. Common examples include:

  • Safety testing where one run refuses a harmful request and another run answers it more permissively.
  • Red-team analysis where a single test pass misses an exploitable response pattern that appears in a later repetition.
  • Regression checks where a model update changes output stability even though the average answer still looks acceptable.
  • Benchmarking workflows where analysts compare pass/fail rates across multiple runs instead of trusting one execution.
  • Operational monitoring where teams look for changing alert quality, citation quality, or classification consistency over time.

A practical tradeoff is that repeated evaluation costs more time and compute, but a single run can give a false sense of reliability. For security-sensitive models, the overhead is often justified because variance can change whether a risk is visible at all.

Security Implications

High run-to-run variance can conceal unsafe behaviour, weaken test confidence, and make controls appear stronger than they are. If an evaluator happens to catch a safe or correct output on one run, they may miss a denial failure, a disclosure, or a policy bypass that emerges only intermittently. That creates a measurement problem as much as a model problem.

For security teams, the main consequence is inconsistent visibility. A model used for classification, triage, summarisation, or policy enforcement may behave differently under the same conditions, which means the organisation cannot assume one test result represents the full operating envelope. The failure mode is often not total collapse, but intermittent unreliability that makes weak controls harder to detect and harder to prove.

A common practitioner observation is that variance often becomes more obvious when prompts are borderline, adversarial, or underspecified. Those are exactly the cases where security review matters most, so repeated evaluation is usually more informative than a single pass.

Domain and Governance Relevance

In AI security governance, run-to-run variance affects how much confidence an organisation can place in testing, monitoring, and acceptance decisions. A model that is acceptable only on some executions creates uncertainty for approval, change control, and incident analysis because the same workload may not produce the same evidence twice.

When the model is used in workflows that influence access, moderation, fraud detection, or security analysis, variance becomes a governance issue rather than a purely technical one. Teams need to know whether instability is an expected property of the system or a sign that the evaluation method is too narrow to support a decision.

For NHI and agentic AI contexts, the relevance is sharper when model output drives actions on behalf of a non-human identity or autonomous agent. If the model’s behaviour is inconsistent, the downstream action chain can become non-deterministic, which complicates ownership, auditability, and safe delegation. That makes repeatability part of trust, not just quality.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

OWASP Agentic AI Top 10 address the attack surface, NIST AI RMF and NIST AI 600-1 set the technical controls, and ISO/IEC 42001:2023 define the regulatory obligations.

FrameworkControl / ReferenceRelevance
NIST AI RMFMEASURE — MeasureRun-to-run variance is a core evaluation stability concern for AI risk measurement.
Recommendation — Measure output stability across repeated runs to detect unreliable model behaviour.
NIST AI 600-1EVAL — EvaluationVariance directly affects the credibility of repeated AI evaluations and benchmarks.
Recommendation — Use repeated evaluations to estimate how often results change under the same input.
ISO/IEC 42001:20238.2 — AI system operationStable operation and evidence-based governance depend on reproducible AI behaviour.
Recommendation — Document acceptable output variability and review it during AI operational governance.
OWASP Agentic AI Top 10A2 — Unreliable or Unexpected Agent BehaviourAgentic systems with variable outputs can act inconsistently across identical runs.
Recommendation — Test for inconsistent agent actions under repeated identical inputs before approval.

Deepen Your Knowledge

Sign up to our weekly newsletter — get 33% off our NHI Foundation Level Course

    NHIMG Editorial Note
    Reviewed and updated by the NHIMG editorial team on September 7, 2026.
    NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org