Phương pháp diễn giải cho mô hình hỏi đáp hình ảnh tiếng Việt
An XAI Method for Vietnamese Visual Question Answering Models
Abstract
Diễn giải trong mô hình hỏi đáp hình ảnh là chủ đề được quan tâm bởi vì tính minh bạch của mô hình sẽ giúp người sử dụng tin tưởng cũng như hiểu về điểm mạnh và hạn chế của mô hình. Các mô hình đa phương thức như hỏi đáp hình ảnh sử dụng cả hai loại dữ liệu hình ảnh và văn bản, do đó gây thách thức cho các kỹ thuật giải thích thường chỉ áp dụng trên một loại dữ liệu. Trong công trình này, chúng tôi đề xuất kết hợp phương pháp giải thích hình ảnh Grad-CAM và phương pháp giải thích văn bản Transformer để phân tích hành vi của mô hình. Các kết quả thu được trên tập dữ liệu hỏi đáp tiếng Việt giúp chúng tôi rút ra các nhận xét như sau: i) Transformer thị giác trích xuất đặc trưng toàn cục trong khi ResNet trích xuất đặc trưng cục bộ; ii) mô hình suy luận không tốt với ảnh có nhiều yếu tố gây nhiễu; iii) phương thức văn bản có ít tác động hơn phương thức hình ảnh; iv) khối PhoBERT có sự thiên vị một số câu hỏi nhất định.
References
A. Weller, Transparency: Motivations and Challenges. Cham: Springer International Publishing, 2019.
R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, and D. Batra, “Grad-cam: Visual explanations
from deep networks via gradient-based localization,” in 2017 IEEE International Conference on Computer Vision (ICCV), pp. 618–626, 2017.
H. Chefer, S. Gur, and L. Wolf, “Transformer interpretability beyond attention visualization,” 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 782–791, 2020.
T. Le, K. Pho, T. Bui, H. T. Nguyen, and M. L. Nguyen, “Object-less vision-language model on visual question classification for blind people,” in Proceedings of the 14th International Conference on Agents and Artificial Intelligence - Volume 3: ICAART, pp. 180–187, SciTePress, Feb. 2022. ICAART 2022, February 3–5, 2022.
A. D. Nguyen, T. Le, and H. T. Nguyen, “Combining multivision embedding in contextual attention for vietnamese visual question answering,” in Image and Video Technology, (Cham), pp. 172–185, Springer International Publishing, 2023.
Y. Lyu, P. P. Liang, Z. Deng, R. Salakhutdinov, and L.- P. Morency, “DIME: Fine-grained interpretations of multimodal models via disentangled local explanations,” 2022. arXiv preprint.
D. H. Park, L. A. Hendricks, Z. Akata, A. Rohrbach, B. Schiele, T. Darrell, and M. Rohrbach, “Multimodal explanations: Justifying decisions and pointing to the evidence,” 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 8779–8788, 2018.
B. Zhou, A. Khosla, A. Lapedriza, A. Oliva, and A. Torralba, “Learning deep features for discriminative localization,” in 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 2921–2929, 2016.
S. Abnar and W. Zuidema, “Quantifying attention flow in transformers,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, (Online), pp. 4190–4197, Association for Computational Linguistics, July 2020.
S. Bach, A. Binder, G. Montavon, F. Klauschen, K.-R. M¨uller, and W. Samek, “On pixel-wise explanations for nonlinear classifier decisions by layer-wise relevance propagation,” PLoS ONE, vol. 10, 2015.
K. Q. Tran, A. T. Nguyen, A. T.-H. Le, and K. V. Nguyen, “Vivqa: Vietnamese visual question answering,” in Pacific Asia Conference on Language, Information and Computation, 2021.
