DaNangVMD: Nhận diện phát âm sai Tiếng Việt

DaNangVMD: Nhận diện phát âm sai tiếng Việt

  • Ket Doan Nguyen
  • Nguyen Anh Tran
  • Van Nam Vo
  • Tran Tien Nguyen
  • Pham Tuyen Le
  • Quoc Vuong Nguyen
  • Huu Nhat Minh Nguyen
##plugins.pubIds.doi.readerDisplayName##: https://doi.org/10.32913/mic-ict-research-vn.v2024.n1.1271
Keywords: Nhận diện phát âm sai, diểu diễn đa phương thức, nhận diện giọng nói tiếng Việt

Abstract

Nhận diện giọng nói tự động, còn được gọi là ASR, đã phát triển mạnh mẽ trong thập kỷ qua và được sử dụng để
nhận diện và dịch giọng nói của con người thành văn bản đọc được một cách tự động. Tuy nhiên, nhận diện giọng
nói tiếng Việt gặp phải những thách thức lớn như lỗi phát âm thường xuyên cũng như sự biến đổi lớn trong giọng
nói tiếng Việt. Trong công trình này, chúng tôi giải quyết thách thức của việc phát hiện lỗi phát âm (MD) trong tiếng
Việt. Là một ngôn ngữ có thanh điệu, tiếng Việt không chỉ dựa trên các phụ âm và nguyên âm mà còn phụ thuộc
vào sự thay đổi trong cao độ hoặc thanh điệu trong quá trình phát âm. Trong bài báo này, chúng tôi đề xuất mô hình
DaNangVMD để phát hiện lỗi phát âm trong giọng nói tiếng Việt dựa trên âm thanh giọng nói và văn bản đúng. Bằng
cách tận dụng biểu diễn đa mô hình dựa trên sự chú ý đa đầu từ các mã hóa của bộ mã hóa ngữ âm và bộ mã hóa
ngôn ngữ, DaNangVMD nhằm cung cấp một giải pháp mạnh mẽ cho việc phát hiện và chẩn đoán lỗi phát âm chính
xác. Qua đánh giá mở rộng, DaNangVMD đề xuất cho thấy hiệu suất vượt trội so với các mô hình PAPL với mức tăng
15% trong điểm F1 và 13% trong độ chính xác.

Author Biographies

Nguyen Anh Tran

Tran Nguyen Anh is pursuing the B.Eng.
degree in Information Technology from
the University of Danang, Vietnam - Korea University of Information and Communication Technology. His research interests include software development, machine
learning and deep learning.
Email: anhtn.21it@vku.udn.vn

Van Nam Vo

Vo Van Nam is pursuing the B.Eng. degree
in Information Technology from the University of Danang, Vietnam - Korea University of Information and Communication
Technology. His research interests include
software development, machine learning
and deep learning.
Email: namvv.21it@vku.udn.vn

Tran Tien Nguyen

Nguyen Tran Tien is pursuing the B.Eng.
degree in Information Technology from
the University of Danang, Vietnam - Korea University of Information and Communication Technology. His research interests include software development, machine
learning and deep learning.
Email: nttien.20it6@vku.udn.vn

Pham Tuyen Le

Le Pham Tuyen received the B.S. degree
in computer science from Ho Chi Minh
City University of Technology, Vietnam, in
2013, and the Ph.D. degree in computer
science and engineering from Kyung Hee
University, South Korea, in 2019. He is
currently an Principal Research Engineer at
AgileSoDA Company and visiting lecturer
at Industrial University of Ho Chi Minh City. His current research
interests include machine learning, reinforcement learning, combinatorial optimization, and robotics.
Email: tuyen_01036033@iuh.edu.vn

Quoc Vuong Nguyen

Nguyen Quoc Vuong receive master degree computer science from the University
of Da Nang, in 2011. His research interests include software development, machine
learing and deep learning.
Email: voung@donga.edu.vn

Huu Nhat Minh Nguyen

Nguyen Huu Nhat Minh (M’20) received
Ph.D. degree in Computer Science and
Engineering from Kyung Hee University,
South Korea, in 2020. He continued PostDoc with Federated Learning and Democratized Learning at Intelligent Networking
lab, Kyung Hee University, South Korea.
He is Deputy Head of Department of Science, Technology, and International Cooperation, and In charge
of Research Program at Digital Science and Technology Institute,
The University of Danang – Vietnam - Korea University of Information and Communication Technology, Vietnam. He received
the best KHU Ph.D. thesis award in engineering in 2020. He
had publications in premier ACM/IEEE journals and conferences.
His research interests include wireless communications, federated
learning, NLP, and computer vision.
Email: nhnminh@vku.udn.vn

References

Wai-Kim Leun. “CNN-RNN-CTC BASED ENDTO-END MISPRONUNCIATION DETECTION AND DIAGNOSIS.” Department of Systems Engineering and Engineering Management, https://www1.se.cuhk.edu.hk/ hccl/publications/pub/ICASSP2019-mdd.pdf. Accessed 22 November 2023.

Kun Li, and Helen Meng. “Mispronunciation Detection and Diagnosis in L2 English Speech Using MultiDistribution Deep Neural Networks.” Human-Computer Communications Laboratory Department of System Engineering and Engineering Management The Chinese University of Hong Kong, Hong Kong SAR, China, 16 June 2023, https://www1.se.cuhk.edu.hk/ hccl/publications/pub/2014 PID3298385 LK.pdf. Accessed 22 November 2023.

Yiqing Feng, et al. “SED-MDD: Towards Sentence Dependent End-To-End Mispronunciation Detection and Diagnosis.” 16 June 2023, https://ieeexplore.ieee.org/document/9052975. Accessed 24 November 2023.

Kaiqi Fu. Jones Lin, Dengfeng Ke, Yanlu Xie, Jinsong Zhang, Binghuai Lin “A Full Text-Dependent End to End Mispronunciation Detection and Diagnosis with Easy Data Augmentation Techniques.” arXiv, 17 April 2021, https://arxiv.org/abs/2104.08428. Accessed 24 November 2023.

Wenxuan Ye, et al. “An Approach to Mispronunciation Detection and Diagnosis with Acoustic, Phonetic and Linguistic (APL) Embeddings.” arXiv, 14 October 2021, https://arxiv.org/abs/2110.07274. Accessed 24 November 2023.

Huu Tuong Tu, et al. “Mispronunciation detection and diagnosis model for tonal language, applied to Vietnamese.” [Dublin, Ireland], 20-24 8 2023, https://www.isca-speech.org/archive/pdfs/interspeech2023/huu23 interspeech.pdf.

Thi Thu Trang Nguyen. HMM-based Vietnamese TextTo-Speech: Prosodic Phrasing Modeling, Corpus Design System Design, and Evaluation, 22 January 2016, https://theses.hal.science/tel-01260884/document. Accessed 24 November 2023.

Maxprotect, https://www.maxprotect.com/. Last accessed 21 Feb. 2024

M. Hammami, Y. Chahir and L. Chen, “WebGuard: Web based adult content detection and filtering system," In

Proceedings IEEE/WIC International Conference on Web Intelligence (WI 2003), Halifax, NS, Canada, 2003, pp. 574-578.

Hu, Weiming, Haiqiang Zuo, Ou Wu, Yunfei Chen, Zhongfei Zhang, and David Suter. “Recognition of adult images, videos, and web page bags." ACM Transactions on Multimedia Computing, Communications, and Applications (TOMM) 7, no. 1 (2011): 1-24.

Sharma, Preeti, Manoj Kumar, and Hitesh Sharma. “Comprehensive analyses of image forgery detection methods from traditional to deep learning approaches: an evaluation." Multimedia Tools and Applications 82, no. 12 (2023): 18117-18150.

Soliman, Mohamed Mostafa, Mohamed Hussein Kamal, Mina Abd El-Massih Nashed, Youssef Mohamed Mostafa, Bassel Safwat Chawky, and Dina Khattab. “Violence recognition from videos using deep learning techniques." In 2019 Ninth International Conference on Intelligent Computing and Information Systems (ICICIS), pp. 80-85. IEEE, 2019.

Human Action Recognition (HAR) Dataset. https://www.kaggle.com/datasets/meetnagadia/humanaction-recognition-har-dataset. Last accessed 21 Feb. 2024

Caruana, Rich. “Multitask learning." Machine learning 28 (1997): 41-75.

M. Ravanelli, T. Parcollet, Y. Bengio, "The PyTorch-Kaldi Speech Recognition Toolkit", https://arxiv.org/abs/1811.07453. Accessed 19 Nov 2018.

Published
2024-05-27