Adversarial malware is a new type of malware that is deliberately devised by attacker for the purpose of evasion from the deep learning-based malware detection system. Although researchers have mainly focused on the devise of novel adversarial malware...
Adversarial malware is a new type of malware that is deliberately devised by attacker for the purpose of evasion from the deep learning-based malware detection system. Although researchers have mainly focused on the devise of novel adversarial malware attacks and defenses, they have been little interested on how the variation of training data set affects the adversarial malware detection capability of the deep leaning-based malware detection system. However, in the sense that the performance of malware detection system based on the deep learning can be varied in accordance with change of the training data set, it is imperative to explore this impact of the training data set variation. To meet this need, we evaluate the impact of training data set variation on the adversarial malware detection system based on the deep learning. We employ MalConv as a basic deep learning-based malware detection system because MalConv is widely used in the field of adversarial malware research. We train MalConv with Dike and Ahn Lab data sets and test these two trained MalConv models with benign CNET and SourceForge software and PartialDos, Gamma, CodeCave, adversarial malware. We present our test results and explain analysis of these results.