Multifamily Malware Models

Basole, Samanvitha; Di Troia, Fabio; Stamp, Mark

Computer Science > Cryptography and Security

arXiv:2207.00620 (cs)

[Submitted on 27 Jun 2022]

Title:Multifamily Malware Models

Authors:Samanvitha Basole, Fabio Di Troia, Mark Stamp

View PDF

Abstract:When training a machine learning model, there is likely to be a tradeoff between accuracy and the diversity of the dataset. Previous research has shown that if we train a model to detect one specific malware family, we generally obtain stronger results as compared to a case where we train a single model on multiple diverse families. However, during the detection phase, it would be more efficient to have a single model that can reliably detect multiple families, rather than having to score each sample against multiple models. In this research, we conduct experiments based on byte $n$-gram features to quantify the relationship between the generality of the training dataset and the accuracy of the corresponding machine learning models, all within the context of the malware detection problem. We find that neighborhood-based algorithms generalize surprisingly well, far outperforming the other machine learning techniques considered.

Subjects:	Cryptography and Security (cs.CR); Machine Learning (cs.LG)
Cite as:	arXiv:2207.00620 [cs.CR]
	(or arXiv:2207.00620v1 [cs.CR] for this version)
	https://doi.org/10.48550/arXiv.2207.00620

Submission history

From: Mark Stamp [view email]
[v1] Mon, 27 Jun 2022 13:06:31 UTC (102 KB)

Computer Science > Cryptography and Security

Title:Multifamily Malware Models

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Cryptography and Security

Title:Multifamily Malware Models

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators