Standard generative adversarial networks for data augmentation in imbalanced intrusion detection system datasets
https://doi.org/10.32362/2500-316X-2026-14-3-7-23
EDN: CZKEAI
Abstract
Objectives. Class imbalance in intrusion detection system (IDS) datasets poses challenges for achieving balanced detection performance. Setting out to evaluate the quality and utility of GAN-generated samples for improving the generalization ability of IDS, this research presents a standard generative adversarial network (GAN) framework as a means of generating synthetic network traffic data for augmenting IDS training datasets. The main focus of the study is an assessment of whether standard GANs can produce realistic synthetic traffic that is not based on targeted generation of minority attack classes.
Methods. When implemented on the NSL-KDD, CIC-IDS2017, and CIC-IDS2018 collections, the GAN displayed precision by mimicking real network traffic distribution. This can be confirmed by inspecting the histograms of different features between flow durations and byte counts and packet rates.
Results. As well as providing stable learning, the presented framework retains diverse sample generation and generates real synthetic data examples. A Random Forest trained on real data achieved 99.86% on the CIC-IDS2017 dataset. This high level of performance was maintained using the GAN-generated synthetic data to confirm the quality of synthetic traffic generation as a tool of overall data augmentation. In contrast, a conventional GAN produces samples based on the total data distribution without focusing on particular attack type, i.e., minority attacks (user-toroot (U2R) and remote-to-local (R2L)) are not explicitly solved.
Conclusions. The paper has shown that conventional GANs have the capability to produce verifiable synthetic network traffic that does not deteriorate overall classifier performance, which validates proof-of-concept of GAN-based data augmentation in IDS. Nonetheless, the conventional GAN is not concerned with minority attack generation, where it produces samples based on the general distribution but without controlling the class. The crucial limitations of the presented methodology are that it is computationally complex and cannot target underrepresented types of attacks (U2R, R2L). Further improvements in conditional GANs should be performed in the future to make them capable of creating class-specific generation and removing class disparity directly in IDS datasets.
About the Authors
Z. ArafatIraq
Zaid Arafat, Assistant Lecturer, Department of Cybersecurity
Scopus Author ID 57963547500
Kerbala, 56001
Competing Interests:
The authors declare no conflicts of interest.
O. V. Yudina
Russian Federation
Olga V. Yudina, Cand.Sci. (Eng.), Associate Professor, Department of Mathematics and Computer Software
5, Lunacharskogo pr., Cherepovets, 162600
Competing Interests:
The authors declare no conflicts of interest.
Z. A. Abdulazeez
Russian Federation
Zainab A. Abdulazeez, Assistant Lecturer, College of Education for Human Sciences
Scopus Author ID 57220186609
Kerbala, 56001
Competing Interests:
The authors declare no conflicts of interest.
References
1. Al-Ajlan M., Ykhlef M. A Review of Generative Adversarial Networks for Intrusion Detection Systems: Advances, Challenges, and Future Directions. Comput. Mater. Contin. 2024;81(2):2053–2076. https://doi.org/10.32604/cmc.2024.055891
2. Arnob A.K.B., Chowdhury R.R., Chaiti N.A., Saha S., Roy A. A comprehensive systematic review of intrusion detection systems: emerging techniques, challenges, and future research directions. J. Edge Comput. 2025;4(1):73–104. https://doi.org/10.55056/jec.885
3. Arifin M.M., Ahmed M.S., Ghosh T.K., Udoy I.A., Zhuang J., Yeh J. A Survey on the Application of Generative Adversarial Networks in Cybersecurity: Prospective. Direction and Open Research Scopes. arXiv preprint arXiv:2407.08839 [cs.CR], 2024. https://doi.org/10.48550/arXiv.2407.08839
4. Kumar V., Sinha D. Synthetic attack data generation model applying generative adversarial network for intrusion detection. Comput. Secur. 2023;125:103054. https://doi.org/10.1016/j.cose.2022.103054
5. Zhao X., Fok K.W., Thing V.L.L. Enhancing Network Intrusion Detection Performance Using Generative Adversarial Networks. Comput. Secur. 2024;145:04005. https://doi.org/10.1016/j.cose.2024.104005
6. Dunmore A., Jang-Jaccard J., Sabrina F., Kwak J. A Comprehensive Survey of Generative Adversarial Networks (GANs) in Cybersecurity Intrusion Detection. IEEE Access. 2023;11:76071–76094. https://doi.org/10.1109/ACCESS.2023.3296707
7. Sabuhi M., Zhou M., Bezemer C.-P., Musilek P. Applications of Generative Adversarial Networks in Anomaly Detection: A Systematic Literature Review. IEEE Access. 2021;9:161003–161029. https://doi.org/10.1109/ACCESS.2021.3131949
8. Zhang S., Xie X., Xu Y. A Brute-Force Black-Box Method to Attack Machine Learning-Based Systems in Cybersecurity. IEEE Access. 2020;8:128250–128263. https://doi.org/10.1109/ACCESS.2020.3008433
9. Shao M., Liu S., Wang R., Zhang G. An Adversarial sample defense method based on multi-scale GAN. Int. J. Mach. Learn. Cybern. 2021;12(2):3437–3447. https://doi.org/10.1007/s13042-021-01374-w
10. Lee J., Park K. GAN-based imbalanced data intrusion detection system. Pers. Ubiquitous Comput. 2021;25(1):121–128. https://doi.org/10.1007/s00779-019-01332-y
11. Lim W., Yong K.S.C., Lau B.T., Tan C.C.L. Future of generative adversarial networks (GAN) for anomaly detection in network security: A review. Comput. Secur. 2024;139:103733. https://doi.org/10.1016/j.cose.2024.103733
12. Shahriar M.H., Haque N.I., Rahman M.A., Alonso M. Jr. G-IDS: Generative Adversarial Networks Assisted Intrusion Detection System. arXiv preprint arXiv:2006.00676 [cs.CR], 2020. https://doi.org/10.48550/arXiv.2006.00676
13. Alotaibi A., Rassam M.A. Adversarial machine learning attacks against intrusion detection systems: A survey on strategies and defense. Future Internet. 2023;15(2):62. https://doi.org/10.3390/fi15020062
14. Achuthan K., Ramanathan S., Srinivas S., Raman R. Advancing cybersecurity and privacy with artificial intelligence: current trends and future research dire ctions. Front. Big Data. 2024;7:1497535. https://doi.org/10.3389/fdata.2024.1497535
15. Almasre M., Subahi A. Create a Realistic IoT Dataset Using Conditional Generative Adversarial Network. J. Sens. Actuator Netw. 2024;13(5):62. https://doi.org/10.3390/jsan13050062
16. Yilmaz I., Masum R., Siraj A. Addressing imbalanced data problem with generative adversarial network for intrusion detection. In: 2020 IEEE 21st International Conference on Information Reuse and Integration for Data Science (IRI). IEEE. 2020. P. 25–30. https://doi.org/10.1109/IRI49571.2020.00012
17. Huang S., Lei K. IGAN-IDS: An imbalanced generative adversarial network towards intrusion detection system in ad-hoc networks. Ad Hoc Netw. 2020;105:102177. https://doi.org/10.1016/j.adhoc.2020.102177
18. Bhattacharya S., Somayaji S., Reddy P.K., et al. A novel PCA-firefly based XGBoost classification model for intrusion detection in networks using GPU. Electronics. 2020;9(2):219. https://doi.org/10.3390/electronics9020219
19. Talukder M.A., Uddin M.A., Hasan K.F., et al. Machine learning-based network intrusion detection for big and imbalanced data using oversampling, stacking feature embedding and feature extraction. J. Big Data. 2024;11(1):33. https://doi.org/10.1186/s40537-024-00886-w
20. Biswas H., Kumar M.M., Kumar P. Intrusion Detection in OT/SCADA Cyber Security and Tabular Generative Adversarial Networks. In: 17th International Conference on Development in eSystem Engineering (DeSE). IEEE. 2024. P. 1–6. https://doi.org/10.1109/DeSE63988.2024.10911912
21. Andresini G., Appice A., De Rose L., Malerba D. GAN augmentation to deal with imbalance in imaging-based intrusion detection. Future Gener. Comput. Syst. 2021;123:108–127. https://doi.org/10.1016/j.future.2021.04.017
22. Al Olaimat M., Lee D., Kim Y., Kim J., Kim J. A learning-based data augmentation for network anomaly detection. In: 2020 29th International Conference on Computer Communications and Networks (ICCCN). IEEE. 2020. P. 1–10. https://doi.org/10.1109/ICCCN49398.2020.9209598
Supplementary files
|
|
1. Generative adversarial network training loss curves for CIC-IDS2017 over 5000 epochs | |
| Subject | ||
| Type | Исследовательские инструменты | |
View
(42KB)
|
Indexing metadata ▾ | |
- It was demonstrated that conventional generative adversarial networks (GANs) have the capability to produce verifiable synthetic network traffic that does not deteriorate overall classifier performance, which validates proof-of-concept of GAN-based data augmentation in intrusion detection system.
- It was shown that the conventional GAN is not concerned with minority attack generation, where it produces samples based on the general distribution but without controlling the class.
Review
For citations:
Arafat Z., Yudina O.V., Abdulazeez Z. Standard generative adversarial networks for data augmentation in imbalanced intrusion detection system datasets. Russian Technological Journal. 2026;14(3):7-23. https://doi.org/10.32362/2500-316X-2026-14-3-7-23. EDN: CZKEAI
JATS XML


























