The assumption that conversion sanitises the model breaks first. If malicious logic is embedded in the graph, ONNX or TensorRT packaging can preserve the trigger path, so the artefact remains operationally dangerous even when it appears to be a normal deployment file.
What actually breaks when conversion does not strip the malicious behaviour?
Format conversion only changes the container, not necessarily the semantics of the model. If the backdoor is embedded in weights, graph structure, or operator patterns, the converted artefact can still respond to the same trigger condition. The practical failure is that teams may trust the new file format as if it were a clean rebuild, when it may still carry the original activation path.
That matters because deployment formats like ONNX or TensorRT are often treated as delivery artefacts, not inspection targets. When conversion preserves the trigger path, the model may remain capable of producing the attacker’s intended output even though the packaging, runtime, and serialization changed around it. The security boundary is the provenance and behaviour of the model, not the file extension.
In other words, conversion can preserve dangerous functionality across ecosystems. A model may pass through export, optimisation, quantisation, or runtime compilation and still retain the same hidden condition that activates the backdoor. The operational risk is not limited to “the original framework is compromised”, it extends to any downstream artefact that faithfully carries the malicious logic forward.
Where the security assumption fails in the conversion pipeline
The failed assumption is that transformation equals sanitisation. In practice, converters are designed to preserve executable behaviour, which is exactly why a backdoor can survive. If the malicious behaviour is represented as a functional path in the model graph rather than as an obvious external script, normal packaging may leave it intact.
That creates a trust gap between source review and deployed behaviour. Engineers may validate that the model exports successfully, but successful conversion only proves compatibility, not innocence. The right question is whether the conversion process removes or alters the triggerable computation, not whether the artefact can be loaded.
For practitioners, the main distinction is between file-format change and behavioural change. A safe conversion should either be accompanied by provenance controls and behavioural testing, or treated as a potentially lossy but untrusted transformation. If the converted model still produces the same trigger-dependent output, the backdoor has not been removed, only repackaged.
Why this is a supply-chain and runtime integrity problem, not just a model-format issue
This failure sits in the AI supply chain because the model that enters deployment is no longer the same artefact that was originally reviewed. The conversion step can become a handoff point where malicious logic survives while the surrounding packaging looks legitimate. That is especially dangerous when teams rely on the converted file as evidence of cleanliness rather than rechecking the model’s behaviour.
The other issue is runtime integrity. If a converted artefact is accepted into production, the trigger may be reachable through ordinary inference requests, which means the backdoor remains operational rather than theoretical. A downstream consumer does not need to know the original training process for the danger to persist.
That is why model provenance, integrity verification, and post-conversion testing are the important controls here. The relevant question is whether the transformation preserved only intended behaviour or also preserved an attacker-controlled branch. If the latter survives, the deployment pipeline has successfully propagated the compromise.
Risk and Threat Considerations
When conversion preserves a backdoor, the main risk is false reassurance: teams may believe they have cleaned or normalised the model when they have only changed its packaging. That can leave a malicious trigger active in production, with the danger hidden behind a trusted deployment format.
Failure mechanism: The conversion pipeline preserves the model’s malicious computation path, so the trigger still activates after export, optimisation, or compilation.
Impact: The artefact can remain exploitable in production, allowing the attacker’s behaviour to survive review, propagate across environments, and evade control assumptions tied to the original file format.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
CSA Cloud Controls Matrix, NIST AI RMF, NIST SP 800-53 Rev 5 and SLSA set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| CSA Cloud Controls Matrix | AIS — AI System Security | Backdoored model conversion is an AI system integrity and supply-chain control issue. |
| Recommendation — Verify model provenance and integrity before promoting converted artefacts. | ||
| NIST AI RMF | GV.1 — Map, Measure, and Manage AI Risks | The question is about AI risk created by preserved malicious behaviour after conversion. |
| Recommendation — Measure post-conversion behaviour and manage residual model risk before deployment. | ||
| NIST SP 800-53 Rev 5 | CM-5 — Access Restrictions for Change | Conversion is a controlled change that can preserve unsafe functionality if unmanaged. |
| Recommendation — Restrict and review model conversion changes before release. | ||
| SLSA | Supply Chain Integrity | The issue is artifact integrity across the AI/model supply chain. |
| Recommendation — Require provenance and integrity evidence for converted model artifacts. | ||
Practitioner Guidance
What to verify: Treat conversion as a behavioural change only when you have evidence that the trigger path was removed, not merely re-encoded. Compare pre- and post-conversion outputs on known trigger inputs, and check whether the converted graph still contains suspicious branches, constants, or activation patterns.
Common mistake: Trusting a successful export or a “standard” runtime format as proof of safety. A converted model can still be malicious, so packaging compatibility should never be confused with model trustworthiness.
What good looks like: The deployment pipeline requires provenance checks, integrity validation, and behavioural tests after every material transformation. If the model changes format, the security review changes with it.
Practitioner takeaway: Conversion is not a sanitiser, it is a transformation step that can faithfully preserve compromise unless you prove otherwise.
Related resources from NHI Mgmt Group
Deepen Your Knowledge
Free weekly newsletter
Subscribe to the NHI & AI Identity Journal
The latest on NHI and Agentic AI security – articles, research, breaches, news and events every week.
Bonus 33% off our NHI Course when you subscribe.
Reviewed and updated by the NHIMG editorial team on October 11, 2026.
NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org