Phát hiện lỗ hổng phần mềm nhờ học sâu

  • Bùi Văn Công
  • Vũ Thảo Nguyên
  • Vũ Đức Minh
  • Nguyễn Phương Lan
##plugins.pubIds.doi.readerDisplayName##: https://doi.org/10.31276/CNTT.2024.3026
Keywords: học sâu, học tương phản, lỗ hổng, phần mềm

Abstract

Việc phát triển các dự án phần mềm thành công luôn là mối quan tâm hàng đầu của các tổ chức, doanh nghiệp. Trong đó, việc đảm bảo chất lượng của phần mềm là ưu tiên cao nhất trong toàn bộ quá trình phát triển cũng như vận hành của phần mềm. Bài báo đề cập việc phát hiện các lỗ hổng của mã nguồn và tập trung vào phân tích cú pháp và ngữ nghĩa của các câu lệnh trong mã nguồn. Mô hình phát hiện lỗ hổng mã nguồn được thực hiện theo quy trình: (i) biểu diễn cú pháp và ngữ nghĩa; (ii) trích xuất đặc trưng của mã nguồn; (iii) cân bằng dữ liệu; (iv) phân loại mã nguồn. Đầu ra của mô hình cho biết mã nguồn đó hoặc bình thường, hoặc có lỗ hổng. Mô hình sử dụng bộ dữ liệu SARD để huấn luyện mô hình học sâu. Mạng học sâu sử dụng mô hình BERT (Bidirectional Encoder Representations from Transformers), mô hình Word2Vec kết hợp LSTM (Long Short-Term Memory) và mô hình Word2Vec với BiLSTM (Bidirectional Long Short-Term Memory) theo ba kịch bản. Kết quả phân loại được đưa qua một hàm softmax để trả về một vector chứa dự đoán xác suất xảy ra của từng loại lỗ hổng. Với tỷ lệ phát hiện lỗ hổng mã nguồn cho kết quả chính xác lên đến 82,63%, tương ứng với tỷ lệ bỏ sót chỉ còn 17,37%. Đây là một kết quả có thể chấp nhận được, chứng minh hiệu quả vượt trội đối với bài toán phát hiện lỗ hổng mã nguồn.

Author Biographies

Bùi Văn Công

Khoa Công nghệ Thông tin, Trường Đại học Kinh tế - Kỹ thuật Công nghiệp, 456 Minh Khai, phường Vĩnh Tuy, Hà Nội, Việt Nam

Vũ Thảo Nguyên

Trường Đại học Công nghệ Nanyang, 50 Đại lộ Nanyang, Singapore

Vũ Đức Minh

Trường THPT Cầu Giấy, 8/118 Nguyễn Khánh Toàn, phường Cầu Giấy, Hà Nội, Việt Nam

Nguyễn Phương Lan

Trường Quốc tế Nhật Bản, 84A Nguyễn Thanh Bình, phường Hà Đông, Hà Nội, Việt Nam

References

[1] B. Chernis, R. Verma (2018), “Machine learning methods for software vulnerability detection,” Proceedings of The Fourth ACM International Workshop on Security and Privacy Analytics, pp.31-39, DOI: 10.1145/3180445.3180453.
[2] B. Liu, W. Guan, C. Yang, et al. (2023), “Transformer and graph convolutional network for text classification”, International Journal of Computational Intelligence Systems, 16, DOI: 10.1007/s44196-023-00337-z.
[3] J. Devlin, M.W. Chang, K. Lee, et al. (2019), “Bert: Pre-training of deep bidirectional transformers for language understanding”, The 2019 Conference of The North American Chapter of The Association for Computational Linguistics: Human Language Technologies, 1, pp.4171-4186, DOI: 10.18653/v1/N19-1423.
[4] D.X. Cho, H.M. Dao, M.T. Cong, et al. (2023), “A novel approach for software vulnerability detection based on intelligent cognitive computing”, The Journal of Supercomputing, 79, pp.17042-17078, DOI: 10.1007/ s11227-023-05282-4.
[5] G. Lin, S. Wen, Q.L. Han, et al. (2020), “Software vulnerability detection using deep neural networks: A survey”, Proceedings of The IEEE, 108(10), pp.1825-1848, DOI: 10.1109/JPROC.2020.2993293.
[6] P. Zeng, G. Lin, L. Pan, et al. (2020), “Software vulnerability analysis and discovery using deep learning techniques: A survey”, IEEE Access, 8, pp.197158-197172, DOI: 10.1109/ACCESS.2020.3034766.
[7] H. Wang, G. Ye, Z. Tang, et al. (2021), “Combining graph-based learning with automated data collection for code vulnerability detection”, IEEE Transactions on Information Forensics and Security, 16, pp.1943-1958, DOI: 10.1109/ TIFS.2020.3044773.
[8] X. Li, L. Wang, Y. Xin, et al. (2021), “Automated software vulnerability detection based on hybrid neural network”, Applied Sciences, 11(7), DOI: 10.3390/ app11073201.
[9] H. Wei, M. Li (2017), “Supervised deep features for software functional clone detection by exploiting lexical and syntactical information in source code”, Proceedings of The TwentySixth International Joint Conference on Artificial Intelligence, pp.3034-3040.
[10] G. Siewruk, W. Mazurczyk (2021), “Contextaware software vulnerability classification using machine learning”, IEEE Access, 9, pp.88852-88867, DOI: 10.1109/ ACCESS.2021.3075385.
[11] X. Li, L. Wang, Y. Xin, et al. (2020), “Automated vulnerability detection in source code using minimum intermediate representation learning”, Appl. Sci., 10, DOI: 10.3390/app10051692.
[12] W. Zheng, J. Gao, X. Wu, et al. (2020), “The impact factors on the performance of machine learningbased vulnerability detection: A comparative study”, The Journal of Systems & Software, 168(7), DOI: 10.1016/j. jss.2020.110659.
[13] R.L. Russell, L. Kim, L.H. Hamilton, et al. (2018), “Automated vulnerability detection in source code using deep representation learning”, 2018 17th IEEE International Conference on Machine Learning and Applications (ICMLA), pp.757-762, DOI: 10.1109/ICMLA.2018.00120.
[14] P. Haridas, G. Chennupati, N. Santhi, et al. (2020), “Code characterization with graph convolutions and capsule networks”, IEEE Access, 8, pp.136307-136315, DOI: 10.1109/ACCESS.2020.3011909.
[15] Z. Li, D. Zou, J. Tang, et al. (2019), “A comparative study of deep learning-based vulnerability detection system”, IEEE Access, 7, pp.103184-103197, DOI: 10.1109/ ACCESS.2019.2930578.
[16] F. Yamaguchi, M. Lottmann, K. Rieck (2012), “Generalized vulnerability extrapolation using abstract syntax trees”, Annual Computer Security Applications Conference, 28, pp.358-368.
[17] H. Gascon, F. Yamaguchi, D. Arp, et al. (2013), “Structural detection of android malware using embedded call graphs”, ACM Workshop on Artificial Intelligence and Security, pp.45-54.
[18] K. Ferrante, J. OttensteinJoe, D. Warren (1989), “The program dependence graph and its use in optimization”, ACM Transactions on Programming Languages and Systems, 9(3), pp.319-349.
[19] D.X. Cho (2023), “A new approach to software vulnerability detection based on CPG analysis”, Cogent Engineering, 10(1), DOI: 10.1080/23311916.2023.2221962.
[20] R. Krishna, Y. Ding, B. Ray (2022), “Deep learning based vulnerability detection: Are we there yet?”, IEEE Transactions on Software Engineering, 48(9), pp.3280-3296, DOI: 10.1109/TSE.2021.3087402.
[21] F. Yamaguchi, N. Golde, D. Arp, et al. (2014), “Modeling and discovering vulnerabilities with code property graphs”, IEEE Symposium on Security and Privacy, pp.590-604, DOI: 10.1109/SP.2014.44.
[22] N.V. Chawla, K.W. Bowyer, L.O. Hall, et al. (2002), “SMOTE: Synthetic minority over-sampling technique”, Journal of Artificial Intelligence Research, 16, pp.321-357.
[23] E. Hoffer, N. Ailon (2015), “Deep metric learning using triplet network”, International Workshop on SimilarityBased Pattern Recognition, pp.84-92.
[24] Z. Li, D. Zou, S. Xu, et al. (2018), “VulDeePecker: A deep learning-based system for vulnerability detection”, arXiv, DOI: 10.14722/ndss.2018.23158.
[25] Z. Li, D. Zou, S. Xu, et al. (2018), “SySeVR: A framework for using deep learning to detect software vulnerabilities”, IEEE Transactions on Dependable and Secure Computing, 19(4), DOI: 10.1109/TDSC.2021.3051525.
Published
2025-12-20