MALIGN: Explainable Static Raw-byte Based Malware Family Classification using Sequence Alignment

Saha, Shoumik; Afroz, Sadia; Rahman, Atif

doi:10.1016/j.cose.2024.103714

Computer Science > Cryptography and Security

arXiv:2111.14185 (cs)

[Submitted on 28 Nov 2021 (v1), last revised 12 Jan 2024 (this version, v3)]

Title:MALIGN: Explainable Static Raw-byte Based Malware Family Classification using Sequence Alignment

Authors:Shoumik Saha, Sadia Afroz, Atif Rahman

View PDF HTML (experimental)

Abstract:For a long time, malware classification and analysis have been an arms-race between antivirus systems and malware authors. Though static analysis is vulnerable to evasion techniques, it is still popular as the first line of defense in antivirus systems. But most of the static analyzers failed to gain the trust of practitioners due to their black-box nature. We propose MAlign, a novel static malware family classification approach inspired by genome sequence alignment that can not only classify malware families but can also provide explanations for its decision. MAlign encodes raw bytes using nucleotides and adopts genome sequence alignment approaches to create a signature of a malware family based on the conserved code segments in that family, without any human labor or expertise. We evaluate MAlign on two malware datasets, and it outperforms other state-of-the-art machine learning based malware classifiers (by 4.49% - 0.07%), especially on small datasets (by 19.48% - 1.2%). Furthermore, we explain the generated signatures by MAlign on different malware families illustrating the kinds of insights it can provide to analysts, and show its efficacy as an analysis tool. Additionally, we evaluate its theoretical and empirical robustness against some common attacks. In this paper, we approach static malware analysis from a unique perspective, aiming to strike a delicate balance among performance, interpretability, and robustness.

Subjects:	Cryptography and Security (cs.CR)
Cite as:	arXiv:2111.14185 [cs.CR]
	(or arXiv:2111.14185v3 [cs.CR] for this version)
	https://doi.org/10.48550/arXiv.2111.14185
Related DOI:	https://doi.org/10.1016/j.cose.2024.103714

Submission history

From: Shoumik Saha [view email]
[v1] Sun, 28 Nov 2021 15:57:28 UTC (842 KB)
[v2] Sun, 20 Aug 2023 13:25:24 UTC (2,176 KB)
[v3] Fri, 12 Jan 2024 21:14:41 UTC (1,788 KB)

Computer Science > Cryptography and Security

Title:MALIGN: Explainable Static Raw-byte Based Malware Family Classification using Sequence Alignment

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Cryptography and Security

Title:MALIGN: Explainable Static Raw-byte Based Malware Family Classification using Sequence Alignment

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators