Code similarity shows that samples may be related at the technical level, while attribution assigns likely authorship, sponsorship, or group identity. Two samples can share much of their code and still differ in operator, campaign, or purpose. Defenders should use similarity to organize analysis, then use infrastructure, targeting, and behavior to support attribution.
How to Separate Technical Resemblance from Attribution
Malware code similarity is a technical comparison. It asks whether two samples share functions, routines, libraries, packing choices, or other implementation traits that suggest a common lineage or reuse. Attribution is a broader judgment. It asks who likely built, deployed, funded, or operated the malware, and that conclusion usually depends on evidence beyond the code itself.
Similarity can be useful even when attribution is weak. Analysts may use it to cluster samples, find variants, identify shared tooling, or connect a campaign to prior malware families. But similar code does not prove the same operator, and dissimilar code does not rule out the same actor if tooling, builders, or suppliers changed.
The practical distinction is scope: similarity describes the artifact, while attribution describes the adversary. That is why two samples can be technically related yet still belong to different campaigns, different operators, or different goals. Attribution becomes stronger only when code findings are corroborated by infrastructure, timing, targeting, language, operational errors, and post-compromise behaviour.
Why Similarity Helps Analysis but Rarely Settles Identity
Code similarity is best treated as an investigative signal, not a conclusion. It helps defenders prioritise reverse engineering, map shared components across samples, and separate reused framework code from genuinely novel payload logic. In practice, this is often enough to say “these samples are related,” but not enough to say “these samples came from the same group.”
Attribution demands a higher evidentiary bar because malware ecosystems are noisy. Builders can be reused, code can be copied, and signatures can be intentionally borrowed or altered. In a professional workflow, similarity supports triage and family grouping, while attribution requires a broader pattern that stands up across multiple sources of evidence.
That is why attribution is inherently probabilistic. It is usually expressed as “likely,” “possible,” or “high confidence,” not as a perfect identification. The more the conclusion depends on code alone, the more fragile it is.
What Changes the Attribution Picture
When defenders move from similarity to attribution, the deciding factors shift toward operational context. Infrastructure reuse, command-and-control patterns, target selection, victim geography, time-of-day behaviour, build artefacts, and shared post-exploitation steps often carry more weight than identical source fragments. Even then, analysts should treat any single indicator as supporting evidence rather than proof.
In supply-chain and credential-theft cases, the code can be almost secondary to what the malware enables. For example, the Shai Hulud npm malware campaign is interesting not just for what the package code did, but for how it exposed secrets and enabled broader abuse. Likewise, the CircleCI Breach illustrates why operator behaviour and downstream access matter more than a narrow code comparison when the real question is who controlled the compromise.
A useful working rule is to use similarity to connect samples, then use operational evidence to test whether those samples belong to the same actor, the same campaign, or merely the same malware ecosystem. That keeps the analysis disciplined and reduces the risk of overclaiming authorship from code overlap alone.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
MITRE ATT&CK and OWASP Non-Human Identity Top 10 address the attack and risk surface, while CIS Controls v8, NIST CSF 2.0 and NIST SP 800-53 Rev 5 set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| MITRE ATT&CK | T1583 — Acquire Infrastructure | Infrastructure reuse and campaign patterns help distinguish attribution from code similarity. |
| Recommendation — Map infrastructure reuse to T1583 and correlate it with observed malware campaigns. | ||
| CIS Controls v8 | CIS-8 — Audit Log Management | Attribution relies on telemetry, logs, and traceable activity across hosts and services. |
| Recommendation — Retain and review logs that preserve attacker activity, infrastructure, and campaign timelines. | ||
| OWASP Non-Human Identity Top 10 | NHI-02 — Secret Leakage | Malware attribution often depends on access paths created by exposed secrets and tokens. |
| Recommendation — Rotate exposed secrets quickly and trace their use across the affected environment. | ||
| NIST CSF 2.0 | DE.CM-01 — Monitoring for Suspicious Activity | Code similarity is only the start; attribution needs monitoring that captures behaviour beyond the sample. |
| Recommendation — Use continuous monitoring to collect behavioural evidence that can support or weaken attribution. | ||
| NIST SP 800-53 Rev 5 | AU-6 — Audit Review, Analysis, and Reporting | Analysts need audit data to corroborate behaviour, timelines, and likely operator identity. |
| Recommendation — Analyze audit records to connect malware activity with infrastructure and campaign evidence. | ||
Practitioner Guidance
What to prioritise: Start by clustering samples on technical similarity, then separate family-level reuse from actor-level evidence. If the code match is strong but infrastructure and targeting diverge, treat attribution as weak or unresolved rather than forcing a single explanation.
What to verify: Look for corroborating signals that survive code refactoring, such as infrastructure overlap, victim selection, reuse of operational mistakes, and consistent post-compromise workflows. Those are often more stable attribution indicators than shared source fragments.
Common mistake: Treating “same code” as “same threat actor” is the fastest way to overstate confidence. Similar malware can reflect shared libraries, borrowed builders, or copied tradecraft, so attribution should be written as a confidence statement, not as a binary verdict.
Practitioner takeaway: Similarity tells you where to investigate next, but attribution is only credible when technical resemblance is joined by repeatable operational evidence.
Related resources from NHI Mgmt Group
- What is the difference between payload similarity and code reuse in malware analysis?
- What is the difference between prompt injection risk and identity abuse in agents?
- What is the difference between SAST and DAST for security teams?
- What is the difference between visual similarity and production-ready code quality?
Deepen Your Knowledge
Reviewed and updated by the NHIMG editorial team on September 29, 2026.
NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org