NORMA eResearch @NCI Library

Enhanced Dynamic Malware Detection from API Call Sequences Using Multi-Architecture Deep Learning and Explainable AI

Londhe, Rigved Devendra (2025) Enhanced Dynamic Malware Detection from API Call Sequences Using Multi-Architecture Deep Learning and Explainable AI. Masters thesis, Dublin, National College of Ireland.

[thumbnail of Master of Science]
Preview
PDF (Master of Science)
Download (876kB) | Preview
[thumbnail of Configuration Manual]
Preview
PDF (Configuration Manual)
Download (136kB) | Preview

Abstract

This is because malware is increasingly becoming more complex and cannot be detected by use of the traditional signature approach. This project focuses on how to detect dynamic malware based on the use of sequences of API calls to the system as behavioral features. Here we both construct and test several models of deep learning, such as a Multi-Layer Perceptron (MLP), 1D Convolutional Neural Network (CNN), Bidirectional Long Short-Term Memory (BiLSTM), a combination of a CNN and an LSTM, and a Transformer (DistilBERT) model, on a cluster of Windows API calls sequences dataset. This dataset, Malware 2025, entails 2,569 logged executions (the samples) that have been classified as malicious or benign, and there are up to 50 ordered API calls that denote each sample (shorter sequences are filtered by padding with an End token). After intensive experiments and hyperparameter optimization, we got detection accuracies of about 9495 with the best models, significantly higher than the ~89 accuracy of prior work. The model with CNN scored the highest (~95.1%), being very close to the MLP and CNN-LSTM (~94.5%), whereas the Transformer model showed a lower result (~91%). We used explainable AI methods in order to make the interpretability higher. A baselined Logistic regression of bag-of-API-call features found salient API calls which were strongly linked to malware (e.g. CopyFileA, MoveFileWithProgressW) or benevolent behavior (e.g. NtQueryKey, GetKeyState). These influential features were also emphasized by Shapley additive explanations (SHAP). The results confirm that the convolutional sequence modeling can effectively distinguish malware based on the runtime calls, and the explainability intertwined adds meaningful information to the rationale of the decision making. Consequently, this study presents an improved malware detection method with great accuracy and transparency of the model, which is important when operating a real-world cybersecurity system.

Item Type: Thesis (Masters)
Supervisors:
Name
Email
Heffernan, Niall
UNSPECIFIED
Subjects: Q Science > QA Mathematics > Electronic computers. Computer science
T Technology > T Technology (General) > Information Technology > Electronic computers. Computer science
Q Science > QH Natural history > QH301 Biology > Methods of research. Technique. Experimental biology > Data processing. Bioinformatics > Artificial intelligence
Q Science > Q Science (General) > Self-organizing systems. Conscious automata > Artificial intelligence
Q Science > QA Mathematics > Computer software > Computer Security
T Technology > T Technology (General) > Information Technology > Computer software > Computer Security
Divisions: School of Computing > Master of Science in Cyber Security
Depositing User: Ciara O'Brien
Date Deposited: 18 Aug 2026 16:46
Last Modified: 18 Aug 2026 16:46
URI: https://norma.ncirl.ie/id/eprint/9540

Actions (login required)

View Item View Item