Một mô hình kiến trúc tổng hợp mới trong mô tả ảnh
Abstract
Tự động sinh ra những mô tả về sự vật, đối tượng trực quan trong ảnh được gọi là chú thích hình ảnh. Nó liên quan đến việc sử dụng cả thị giác máy tính và xử lý ngôn ngữ tự nhiên (NLP) để trích xuất các đặc điểm của hình ảnh. Trong nghiên cứu này, chúng tôi phát triển mô hình SPNER-SAN kết hợp giữa Mạng tự chú ý quan hệ nâng cao (ER-SAN), có dạng cấu trúc giống như Transformer, với Mạng đề xuất đồ thị con (SPN) mới được phát triển của chúng tôi. Mục tiêu chính của sự kết hợp này là cải thiện kết quả của quá trình sinh chú thích cho hình ảnh bằng cách tận dụng các ưu điểm của cả hai thành phần. Thực nghiệm trên bộ dữ liệu Coco chứng minh rằng mô hình SPNER-SAN cải thiện đáng kể chất lượng của câu mô tả sinh ra so với với các phương pháp liên quan. Mô hình đạt hiệu suất cao hơn trên cả 5 chỉ số độ đo so sánh như BLEU, METEOR, ROUGR, CIDER, và SPICE. Điều này cho thấy tính hiệu quả của nó trong việc sinh tự động các chú thích rõ ràng và phù hợp với ngữ cảnh.
References
T. Ghandi, H. Pourreza, and H. Mahyar, “Deep learning approaches on image captioning: A review,” ACM Computing Surveys, vol. 56, no. 3, pp. 1–39, 2023.
A. Selivanov, O. Y. Rogov, D. Chesakov, A. Shelmanov, I. Fedulova, and D. V. Dylov, “Medical image captioning via generative pretrained transformers,” Scientific Reports, vol. 13, no. 1, p. 4171, 2023.
M. Shoman, D. Wang, A. Aboah, and M. Abdel-Aty, “Enhancing traffic safety with parallel dense video captioning for end-to-end event analysis,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 7125–7133.
Y. Pan, T. Yao, Y. Li, and T. Mei, “X-linear attention networks for image captioning,” in Proceedings of the
IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 10 971–10 980.
K. Safiya and R. Pandian, “A real-time image captioning framework using computer vision to help the visually impaired,” Multimedia Tools and Applications, vol. 83, no. 20, pp. 59 413–59 438, 2024.
Z. Song, X. Zhou, L. Dong, J. Tan, and L. Guo, “Direction relation transformer for image captioning,” in Proceedings of the 29th ACM International Conference on Multimedia, 2021, pp. 5056–5064.
J. Li, Z. Mao, S. Fang, and H. Li, “Er-san: Enhancedadaptive relation self-attention network for image captioning.” in IJCAI, 2022, pp. 1081–1087.
S. Ren, K. He, R. Girshick, and J. Sun, “Faster r-cnn: Towards real-time object detection with region proposal networks,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 39, no. 6, pp. 1137–1149, 2016.
P. Anderson, X. He, C. Buehler, D. Teney, M. Johnson, S. Gould, and L. Zhang, “Bottom-up and top-down attention for image captioning and visual question answering,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 6077–6086.
H. Wu, Y. Liu, H. Cai, and S. He, “Learning transferable perturbations for image captioning,” ACM Transactions on Multimedia Computing, Communications, and Applications (TOMM), vol. 18, no. 2, pp. 1–18, 2022.
M. Yuan, B.-K. Bao, Z. Tan, and C. Xu, “Adaptive text denoising network for image caption editing,” ACM Transactions on Multimedia Computing, Communications and Applications, vol. 19, no. 1s, pp. 1–18, 2023.
S. Dubey, F. Olimov, M. A. Rafique, J. Kim, and M. Jeon, “Label-attention transformer with geometrically coherent objects for image captioning,” Information Sciences, vol. 623, pp. 812–831, 2023.
T. Yao, Y. Pan, Y. Li, and T. Mei, “Exploring visual relationship for image captioning,” in Proceedings of the
European conference on computer vision (ECCV), 2018, pp. 684–699.
X. Yang, H. Zhang, and J. Cai, “Learning to collocate neural modules for image captioning,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2019, pp. 4250–4260.
Z. Shi, X. Zhou, X. Qiu, and X. Zhu, “Improving image captioning with better use of captions,” arXiv preprint
arXiv:2006.11807, 2020.
A.-A. Liu, Y. Zhai, N. Xu, W. Nie, W. Li, and Y. Zhang, “Region-aware image captioning via interaction learning,”
IEEE Transactions on Circuits and Systems for Video Technology, vol. 32, no. 6, pp. 3685–3696, 2021.
A. Abedi, H. Karshenas, and P. Adibi, “Multi-modal reward for visual relationships-based image captioning,” arXiv preprint arXiv:2303.10766, 2023.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 770–778.
P. Velickovic, G. Cucurull, A. Casanova, A. Romero, P. Lio, Y. Bengio et al., “Graph attention networks,” stat, vol. 1050, no. 20, pp. 10–48 550, 2017.
A. Vaswani, “Attention is all you need,” Advances in Neural Information Processing Systems, 2017.
X. Li, G. Zhang, Y.Wu, X. Li, and Y. Zhang, “Rule of thirdsaware reinforcement learning for image aesthetic cropping,” The Visual Computer, vol. 39, no. 11, pp. 5651–5667, 2023.
M. Freitag and Y. Al-Onaizan, “Beam search strategies for neural machine translation,” arXiv preprint arXiv:1702.01806, 2017.
A. Karpathy and L. Fei-Fei, “Deep visual-semantic alignments for generating image descriptions,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2015, pp. 3128–3137.
J. Li, Z. Mao, H. Li, W. Chen, and Y. Zhang, “Exploring visual relationships via transformer-based graphs for enhanced image captioning,” ACM Transactions on Multimedia Computing, Communications and Applications, vol. 20, no. 5, pp. 1–23, 2024.
