Join our Newsletter — 33% off our NHI Course

Should organisations prioritise adversarial testing or runtime monitoring first?

They need both, but adversarial testing should come first in the promotion path because it establishes whether the model is safe to release at all. Runtime monitoring then catches drift and behavioural anomalies after deployment. If a model is never tested against malicious inputs before release, monitoring alone only tells you when the failure is already live.

Why adversarial testing belongs before release decisions

adversarial testing answers the question that runtime monitoring cannot answer on its own: whether the model behaves safely when deliberately stressed before it is promoted. If malicious prompting, tool abuse, jailbreak attempts, or boundary-pushing inputs can trigger unsafe output or actions in a controlled test, the release decision should be delayed until those failure modes are understood and reduced.

This is why adversarial testing is a promotion gate, not just a quality activity. It forces teams to validate the model against misuse cases, unsafe delegation, and prompt-injection style behaviours while the system is still changeable. Runtime monitoring becomes more valuable once a model is already deployed, because then the concern shifts from release readiness to ongoing drift, abuse, and anomaly detection.

Adversarial testing also helps teams distinguish “works in normal use” from “fails under pressure.” That distinction matters most for systems that can trigger real actions, call tools, or influence downstream decisions. A model that passes only benign evaluation still may be too brittle for release if stress testing reveals a narrow safety margin.

What runtime monitoring is for after deployment

Runtime monitoring is the control that watches live behaviour once the model is in production. It looks for drift, abnormal outputs, policy violations, unusual tool use, emerging prompt patterns, or changes in the operating environment that were not visible during testing. It is a detection and response layer, not a substitute for pre-release validation.

Monitoring becomes essential because production conditions change. User behaviour shifts, integrations expand, prompt patterns evolve, and attackers adapt. Even a model that was adequately tested can become risky later if the surrounding system changes or if new failure paths emerge. The practical value of monitoring is that it catches problems after deployment, but before they become widespread or irreversible.

Good monitoring is specific to the model’s actual blast radius. If the model can generate recommendations only, alerting thresholds can be looser. If it can approve actions, invoke tools, or expose data, monitoring must be tighter and tied to meaningful runtime signals. In other words, the more authority the model has, the more important continuous observation becomes.

How to sequence both controls without overdoing either one

The right sequence is to test first, then monitor, then keep feeding runtime findings back into the next testing cycle. That sequencing reflects different jobs: adversarial testing establishes release readiness, while runtime monitoring manages live operational risk. One tells you whether the system should go out; the other tells you what is happening once it is out.

Teams often make the mistake of overvaluing dashboards before they have a proper adversarial baseline. If you have not tested the model against malicious or edge-case inputs, the monitor may simply document failure as it happens. A better operating model is to use testing to define expected safe behaviour, then use monitoring to detect deviation from that baseline.

For practical governance, release criteria should be explicit, and monitoring should be linked to escalation paths. If adversarial testing finds a repeatable exploit or unsafe tool path, that is a release blocker. If monitoring later shows the same class of issue emerging in production, the response should be containment, rollback, or prompt and policy hardening rather than passive observation.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

MITRE ATLAS and OWASP Agentic AI Top 10 address the attack and risk surface, while NIST AI RMF and NIST CSF 2.0 set the governance and control requirements practitioners need to meet.

Framework Control / Reference Relevance
MITRE ATLAS MITRE ATLAS adversarial AI threat matrix Covers adversarial AI techniques that testing should simulate before release.
Recommendation — Map likely attack paths to ATLAS techniques and test them before promotion.
OWASP Agentic AI Top 10 ASI03 — Identity & Privilege Abuse Directly applies when adversarial testing checks whether the model can misuse delegated privileges or tools.
Recommendation — Test for privilege abuse and block release until unsafe authority paths are closed.
NIST AI RMF GOVERN — Govern Supports pre-release governance decisions and ongoing oversight for AI systems.
Recommendation — Define release gates and monitoring ownership before deployment.
NIST CSF 2.0 DE.CM-01 — Monitoring for Anomalies, Events, and Continuous Security Monitoring Matches the runtime monitoring function of detecting abnormal live behaviour.
PR.DS-01 — Data-at-rest is protected Supports protecting data exposed through model operation and monitoring pipelines.
Recommendation — Instrument production to detect behavioural anomalies and alert on drift. Protect sensitive data paths used by the model and its monitoring stack.

Practitioner Guidance

What to prioritise: Put adversarial testing ahead of any promotion decision, especially when the model can influence tools, workflows, or external decisions. Treat runtime monitoring as mandatory for live operations, but not as evidence that release was safe in the first place.

What to verify: Confirm that your test plan includes malicious input, jailbreak, prompt-injection, tool-abuse, and escalation scenarios that reflect the model’s real authority. Also verify that monitoring has alerts tied to behaviour that matters, not just generic traffic volume or token counts.

Decision rule: If the model has not been tested against plausible abuse cases, do not let monitoring be the only control standing between the model and production. If it has been tested and approved, keep monitoring focused on drift, anomalous actions, and newly emerging attack patterns.

Practitioner takeaway: Adversarial testing is the release gate, runtime monitoring is the live safeguard, and the two only work well when testing defines what “safe” means before production begins.