Corrected runs are training traces where a stronger system or reviewer makes the smallest necessary fixes to a model’s attempted actions. They are useful because they preserve near-valid behavior and expose the exact adjustment the model needs to learn, rather than replacing the model’s output entirely.
What Corrected Runs Are Used For
Corrected runs are most useful when the goal is to preserve a model’s near-correct behavior while teaching the exact boundary where its action should change. They sit between raw attempts and fully replaced labels, making them especially valuable for fine-grained training and review workflows.
Because the correction is minimal, the trace retains useful structure: the original intent, the sequence leading up to the mistake, and the smallest repair that makes the output acceptable. That makes corrected runs a practical teaching signal when you want improvement without erasing the model’s original decision path.
How Corrected Runs Differ From Full Rewrites
A full rewrite replaces the model’s output with a clean target, but a corrected run keeps the attempted action and adjusts only what is necessary. That difference matters because the model is trained on the relationship between the attempt and the correction, not just on the final answer.
This approach is usually better when the error is local, such as a wrong field, a missing step, or a small policy violation. If the original attempt is too far off, a corrected run can become noisy and less instructive than a complete replacement trace.
Why Corrected Runs Improve Learning Signal Quality
Corrected runs preserve the context around the mistake, which helps separate robust behavior from the specific failure. The model can see what almost worked, what was corrected, and how little change was required to move from invalid to valid.
That makes them a strong fit for supervising action-oriented systems where the important question is not only whether the answer is right, but whether the sequence of steps was safe, precise, and recoverable. They are also helpful when reviewers want to standardise judgment across many similar cases without flattening everything into generic labels.
Where Corrected Runs Fit in Training and Review Pipelines
Corrected runs are typically used in datasets or review loops where human or higher-capability oversight can make targeted edits to a candidate trace. The resulting record can then be reused for model improvement, evaluator calibration, or policy tuning.
They are most effective when the correction policy is consistent, because the value of the signal depends on the reviewer making the smallest necessary change rather than introducing a new style or preference. When used well, corrected runs create a compact record of the expected adjustment and reduce ambiguity about what the model should learn next.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
NIST SP 800-53 Rev 5 and OWASP ASVS set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| NIST SP 800-53 Rev 5 | SI-4 — System Monitoring | Corrected runs help surface precise deviations that monitoring and review processes must detect. |
| AU-6 — Audit Review, Analysis, and Reporting | Corrected runs preserve the original attempt and the minimal fix, which supports review and analysis. | |
| Recommendation — Use SI-4 to review corrected traces for repeatable failure patterns and policy deviations. Apply AU-6 to compare attempted actions with reviewer corrections and document recurring errors. | ||
| OWASP ASVS | V15 — Secure Coding and Architecture | Corrected runs are a feedback technique for improving action sequences and implementation quality. |
| Recommendation — Use V15 to encode corrected examples into repeatable secure-behavior training and review. | ||
Practitioner Guidance
What to watch for: Use corrected runs when the failure is local and the surrounding trace is still valuable. If reviewers start making broad edits, the example is no longer a corrected run in practice, it is becoming a rewrite.
Governance implication: Teams should define what counts as a minimal correction so review quality stays consistent across annotators and use cases. That policy choice directly affects dataset fidelity, evaluator agreement, and how well the training signal reflects the intended behavior.
Related resources from NHI Mgmt Group
- Why do corrected training runs improve agent performance more than other supervision signals?
- Who is accountable when an AI agent runs a query on behalf of a user?
- What breaks when MCP runs behind gateways without defined auth propagation?
- How should security teams handle a supply-chain malware event that runs during npm install?