Behavior-Aware Static Malware Detection: Benhmarking Classical and Modern Machine Learning Models on the EMBER Dataset

Authors

  • Intan Permata Sari STMIK Eresha Author
  • Rita Hamdani STMIK Eresha Author

Keywords:

static malware detection, behavior-aware feature engineering, EMBER dataset, explainable machine learning, XGBoost, SHAP analysis, cybersecurity benchmarking, portable executable analysis

Abstract

Malware attacks continue to evolve rapidly in sophistication, scale, and obfuscation capability, posing significant challenges to traditional signature-based cybersecurity systems. The growing diversity of malicious software behaviors has increased the demand for scalable and intelligent malware detection frameworks capable of maintaining strong predictive performance while remaining interpretable and computationally practical. Despite substantial advances in machine learning-based malware detection, many existing studies primarily emphasize classification accuracy while providing limited discussion regarding explainability, computational tradeoffs, and fair benchmarking between classical machine learning approaches and modern deep learning architectures. In particular, the effectiveness of lightweight behavior-aware feature engineering within large-scale static malware analysis remains insufficiently explored. This study presents a comprehensive benchmarking framework for static malware detection using the EMBER dataset. The proposed framework evaluates several classical and modern learning approaches, including XGBoost, Random Forest, Multi-Layer Perceptron (MLP), and TabNet, while incorporating behavior-aware import-based features extracted from Portable Executable (PE) files. In addition, SHAP explainability analysis is employed to investigate feature contribution and model decision behavior. Experimental results demonstrate that optimized classical ensemble-based approaches consistently outperform several modern deep learning models across multiple evaluation metrics. The tuned XGBoost configuration achieved the best overall performance, reaching 94.43% classification accuracy and a ROC-AUC score of 0.9871. SHAP analysis further revealed that entropy characteristics, import statistics, suspicious API indicators, and string-related features were among the most influential contributors to malware classification decisions. This work contributes a large-scale benchmarking framework integrating behavior-aware feature engineering, explainable AI analysis, and computational tradeoff evaluation for practical malware detection systems. Future research may extend this framework through dynamic malware analysis, transformer-based architectures, and adversarially robust detection strategies.

Downloads

Download data is not yet available.

References

Alani, M. M., Mashatan, A., & Miri, A. (2023a). XMal: A lightweight memory-based explainable obfuscated-malware detector. Computers & Security, 133, 103409. https://doi.org/10.1016/j.cose.2023.103409

Alani, M. M., Mashatan, A., & Miri, A. (2023b). XMal: A lightweight memory-based explainable obfuscated-malware detector. Computers & Security, 133, 103409. https://doi.org/10.1016/j.cose.2023.103409

Alenezi, R., & Ludwig, S. A. (2021a). Explainability of Cybersecurity Threats Data Using SHAP. 2021 IEEE Symposium Series on Computational Intelligence (SSCI), 01–10. https://doi.org/10.1109/SSCI50451.2021.9659888

Alenezi, R., & Ludwig, S. A. (2021b). Explainability of Cybersecurity Threats Data Using SHAP. 2021 IEEE Symposium Series on Computational Intelligence (SSCI), 01–10. https://doi.org/10.1109/SSCI50451.2021.9659888

Al-Issa, A. (2005). The Role of English Language Culture in the Omani Language Education System: An Ideological Perspective. Language, Culture and Curriculum, 18(3), 258–270. https://doi.org/10.1080/07908310508668746

Aliyah, I. (2016). The Roles of Traditional Markets as the Main Component of Javanese Culture Urban Space (Case Study: The City of Surakarta, Indonesia). IAFOR Journal of Sustainability, Energy & the Environment, 3(1). https://doi.org/10.22492/ijsee.3.1.06

Alzu’bi, A., Abuarqoub, A., Abdullah, M., Agolah, R. A., & Al Ajlouni, M. (2024). Malware Prediction Using Tabular Deep Learning Models (pp. 379–389). https://doi.org/10.1007/978-3-031-47508-5_30

Ambekar, N. G., Devi, N. N., Thokchom, S., & Yogita. (2025). TabLSTMNet: enhancing android malware classification through integrated attention and explainable AI. Microsystem Technologies, 31(3), 695–713. https://doi.org/10.1007/s00542-024-05615-0

Ashawa, M., Owoh, N., Hosseinzadeh, S., & Osamor, J. (2024). Enhanced Image-Based Malware Classification Using Transformer-Based Convolutional Neural Networks (CNNs). Electronics, 13(20), 4081. https://doi.org/10.3390/electronics13204081

Balasubramanian, K. M., Vasudevan, S. V., Thangavel, S. K., T, G. K., Srinivasan, K., Tibrewal, A., & Vajipayajula, S. (2023). Obfuscated Malware detection using Machine Learning models. 2023 14th International Conference on Computing Communication and Networking Technologies (ICCCNT), 1–8. https://doi.org/10.1109/ICCCNT56998.2023.10307598

Barbero, F., Pendlebury, F., Pierazzi, F., & Cavallaro, L. (2022). Transcending TRANSCEND: Revisiting Malware Classification in the Presence of Concept Drift. 2022 IEEE Symposium on Security and Privacy (SP), 805–823. https://doi.org/10.1109/SP46214.2022.9833659

Darmawan, I., Setiaji, H., Sutriani, L., Supriyadi, A., Rizal, R., & Rahmatulloh, A. (2025). Ensemble Learning for Malware Classification: A Performance Evaluation Using Random Forest and Hist Gradient Boosting. 2025 International Conference on Information and Communication Technology (ICoICT), 1–6. https://doi.org/10.1109/ICoICT66265.2025.11193116

El Neel, L., Copiaco, A., Obaid, W., & Mukhtar, H. (2022). Comparison of Feature Extraction and Classification Techniques of PE Malware. 2022 5th International Conference on Signal Processing and Information Security (ICSPIS), 26–31. https://doi.org/10.1109/ICSPIS57063.2022.10002693

Galen, C., & Steele, R. (2020). Evaluating Performance Maintenance and Deterioration Over Time of Machine Learning-based Malware Detection Models on the EMBER PE Dataset. 2020 Seventh International Conference on Social Networks Analysis, Management and Security (SNAMS), 1–7. https://doi.org/10.1109/SNAMS52053.2020.9336538

Guven, M. (2024). Dynamic Malware Analysis Using a Sandbox Environment, Network Traffic Logs, and Artificial Intelligence. International Journal of Computational and Experimental Science and Engineering, 10(3). https://doi.org/10.22399/ijcesen.460

Hafiz, Md. F. Bin, Khan, N. A., Kamal, Z., Hossain, S., & Barman, S. (2024). A Robust Malware Classification Approach Leveraging Explainable AI. 2024 International Conference on Intelligent Systems for Cybersecurity (ISCS), 1–6. https://doi.org/10.1109/ISCS61804.2024.10581382

Hasan, R., Biswas, B., Samiun, M., Saleh, M. A., Prabha, M., Akter, J., Joya, F. H., & Abdullah, M. (2025). Enhancing malware detection with feature selection and scaling techniques using machine learning models. Scientific Reports, 15(1), 9122. https://doi.org/10.1038/s41598-025-93447-x

Hashim, M. et al. (2024). Enhancing XGBoost Performance in Malware Detection through Chi-Squared Feature Selection. https://www.researchgate.net/publication/386096117

Imran, M. F. (2025). Malware classification using SVM and XGBoost: A study of the influence of features. Journal of Al-Qadisiyah for Computer Science and Mathematics, 17(4). https://doi.org/10.29304/jqcsm.2025.17.42536

Kacem, T., & Tossou, S. (2025). Trandroid: An Android Mobile Threat Detection System Using Transformer Neural Networks. Electronics, 14(6), 1230. https://doi.org/10.3390/electronics14061230

Karat, G., Kannimoola, J. M., Nair, N., Vazhayil, A., G, S. V, & Poornachandran, P. (2024). CNN-LSTM Hybrid Model for Enhanced Malware Analysis and Detection. Procedia Computer Science, 233, 492–503. https://doi.org/10.1016/j.procs.2024.03.239

Kim, H., & Kim, M. (2024). Malware Detection and Classification System Based on CNN-BiLSTM. Electronics, 13(13), 2539. https://doi.org/10.3390/electronics13132539

Lai, T.-H., Tsai, Y.-J., & Liu, C.-L. (2025). Improving the Performance of Static Malware Classification Using Deep Learning Models and Feature Reduction Strategies. Mathematics, 13(23), 3753. https://doi.org/10.3390/math13233753

Ling, X., Wu, L., Zhang, J., Qu, Z., Deng, W., Chen, X., Qian, Y., Wu, C., Ji, S., Luo, T., Wu, J., & Wu, Y. (2023). Adversarial attacks against Windows PE malware detection: A survey of the state-of-the-art. Computers & Security, 128, 103134. https://doi.org/10.1016/j.cose.2023.103134

M., G., & Sethuraman, S. C. (2023). A comprehensive survey on deep learning based malware detection techniques. Computer Science Review, 47, 100529. https://doi.org/10.1016/j.cosrev.2022.100529

Maniriho, P., Mahmood, A. N., & Chowdhury, M. J. M. (2022). A study on malicious software behaviour analysis and detection techniques: Taxonomy, current trends and challenges. Future Generation Computer Systems, 130, 1–18. https://doi.org/10.1016/j.future.2021.11.030

Naseer, M., Ullah, F., Saeed, S., Algarni, F., & Zhao, Y. (2025). Explainable TabNet ensemble model for identification of obfuscated URLs with features selection to ensure secure web browsing. Scientific Reports, 15(1), 9496. https://doi.org/10.1038/s41598-025-93286-w

Nazim, S., Alam, M. M., Rizvi, S. S., Mustapha, J. C., Hussain, S. S., & Suud, M. M. (2025). Advancing malware imagery classification with explainable deep learning: A state-of-the-art approach using SHAP, LIME and Grad-CAM. PLOS One, 20(5), e0318542. https://doi.org/10.1371/journal.pone.0318542

Oyama, Y., Miyashita, T., & Kokubo, H. (2019). Identifying Useful Features for Malware Detection in the Ember Dataset. 2019 Seventh International Symposium on Computing and Networking Workshops (CANDARW), 360–366. https://doi.org/10.1109/CANDARW.2019.00069

Rathod, V., Parekh, C., & Dholariya, D. (2021). AI & ML Based Anamoly Detection and Response Using Ember Dataset. 2021 9th International Conference on Reliability, Infocom Technologies and Optimization (Trends and Future Directions) (ICRITO), 1–5. https://doi.org/10.1109/ICRITO51393.2021.9596451

Şandor, M., Portase, R. M., & Coleşa, A. (2023). Ember Feature Dataset Analysis For Malware Detection. 2023 IEEE 19th International Conference on Intelligent Computer Communication and Processing (ICCP), 203–210. https://doi.org/10.1109/ICCP60212.2023.10398693

Saqib, M., Mahdavifar, S., Fung, B. C. M., & Charland, P. (2024a). A Comprehensive Analysis of Explainable AI for Malware Hunting. ACM Computing Surveys, 56(12), 1–40. https://doi.org/10.1145/3677374

Saqib, M., Mahdavifar, S., Fung, B. C. M., & Charland, P. (2024b). A Comprehensive Analysis of Explainable AI for Malware Hunting. ACM Computing Surveys, 56(12), 1–40. https://doi.org/10.1145/3677374

Tsuchiya, K. (2018). Javanology and the Age of Ranggawarsita: An Introduction to Nineteenth-Century Javanese Culture. In T. Shiraishi (Ed.), Reading Southeast Asia (pp. 75–108). Cornell University Press. https://doi.org/10.7591/9781501718922-005

Ucci, D., Aniello, L., & Baldoni, R. (2019a). Survey of machine learning techniques for malware analysis. Computers & Security, 81, 123–147. https://doi.org/10.1016/j.cose.2018.11.001

Ucci, D., Aniello, L., & Baldoni, R. (2019b). Survey of machine learning techniques for malware analysis. Computers & Security, 81, 123–147. https://doi.org/10.1016/j.cose.2018.11.001

Widodo, S. T. (2020). Norms and teachings in the art of lovemaking of kings in ancient Javanese manuscripts. Rupkatha Journal on Interdisciplinary Studies in Humanities, 12(1), 1–12. https://doi.org/10.21659/rupkatha.v12n1.30

Wijaya, D. A., Djono, D., & Ediyono, S. (2018). Local Knowledge in Joglo Majapahit: Analysis of Local Wisdom Models Gemah Ripah Loh Jinawi in Rural Java. International Journal of Multicultural and Multireligious Understanding, 5(3), 113. https://doi.org/10.18415/ijmmu.v5i3.235

Yousuf, M. I., Anwer, I., Riasat, A., Zia, K. T., & Kim, S. (2023). Windows malware detection based on static analysis with multiple features. PeerJ Computer Science, 9, e1319. https://doi.org/10.7717/peerj-cs.1319

Zhang, Z., Hamadi, H. Al, Damiani, E., Yeun, C. Y., & Taher, F. (2022). Explainable Artificial Intelligence Applications in Cyber Security: State-of-the-Art in Research. IEEE Access, 10, 93104–93139. https://doi.org/10.1109/ACCESS.2022.3204051

Downloads

Published

2026-01-15

Issue

Section

Articles

How to Cite

Behavior-Aware Static Malware Detection: Benhmarking Classical and Modern Machine Learning Models on the EMBER Dataset. (2026). International Journal of Applied Information Systems and Cybersecurity, 1(1), 82-103. https://ejournal.global-scientificjournal.org/ijaisc/article/view/64