Tiếng Anh Phi-3 với Luật pháp: Tinh chỉnh mô hình ngôn ngữ nhỏ trong việc hiểu văn bản luật

  • Hữu Khánh Nguyễn Thai Nguyen University, Thai Nguyen, Vietnam
  • Van Viet Nguyen Thai Nguyen University of Information and Communication Technology, Thai Nguyen, Vietnam
  • Nguyen The Vinh Thai Nguyen University of Information and Communication Technology, Thai Nguyen, Vietnam
  • Nguyen Huu Cong Thai Nguyen University, Thai Nguyen, Vietnam
Keywords: Phi-3, CaseHold, Fine-tuning, QLoRA, Supervised fine-tuning

Abstract

 Nghiên cứu này khám phá ứng dụng mô hình Phi-3 của Microsoft vào việc hiểu văn bản pháp lý. Chúng tôi đã tinh chỉnh Phi-3-mini-4k-instruct, một mô hình ngôn ngữ nhỏ gọn nhưng mạnh mẽ, trên hơn 53.000 câu hỏi trắc nghiệm của CaseHOLD được thiết kế để kiểm tra khả năng nhận dạng vụ án pháp lý. Sử dụng các kỹ thuật thích ứng bậc thấp (LoRA) và thích ứng bậc thấp lượng tử (QLoRA), chúng tôi đã tối ưu hóa quy trình tinh chỉnh để đạt hiệu quả. Kết quả của chúng tôi cho thấy mô hình Phi-3-mini-4k-instruct tinh chỉnh đạt điểm F1 là 76,89, vượt qua các mô hình tiên tiến trước đây về khả năng hiểu văn bản pháp lý. Hiệu suất này đạt được chỉ với 6000 bước đào tạo, làm nổi bật khả năng thích ứng nhanh chóng của Phi-3 với các tác vụ cụ thể theo miền. Mô hình Phi-3-mini-4k-instruct cơ bản, không cần tinh chỉnh, đã đạt điểm F1 là 64,89, vượt trội hơn một số mô hình pháp lý chuyên biệt. Kết quả thực nghiệm của chúng tôi nhấn mạnh tiềm năng của các mô hình ngôn ngữ nhỏ gọn như Phi-3 trong các miền chuyên biệt khi được tinh chỉnh đúng cách, mang lại sự cân bằng giữa quy mô mô hình, hiệu quả đào tạo và hiệu suất. Nghiên cứu này góp phần vào khối lượng công việc ngày càng tăng về việc điều chỉnh các mô hình ngôn ngữ lớn cho các miền cụ thể, đặc biệt là trong lĩnh vực pháp lý.

Author Biographies

Hữu Khánh Nguyễn, Thai Nguyen University, Thai Nguyen, Vietnam

Nguyen Huu Khanh has graduated with a Master’s degree in Computer Science from the University of Information and Communications Technology- Thai Nguyen University since 2022 and is currently a PhD
student here since 2023. His main research interests are Computer Science, Natural Language Processing and Computer Vision.

Van Viet Nguyen, Thai Nguyen University of Information and Communication Technology, Thai Nguyen, Vietnam

Nguyen Van Viet received the Bachelor’s Information Technology at Thai Nguyen University in 2009 and Master’s degree Information Technology at Manuael S. Enverga University Foundation, Lucena City,
Philippines in 2012. He worked as a lecturer at the Faculty of Information Technology, School of Information and Communication Technology, Thai Nguyen University from 2009. Now, he is a researcher at the Thai Nguyen University of Information and Communication Technology, Thai Nguyen, Vietnam. Office address: University of Information and Communication Technology, Thai Nguyen University, Thai Nguyen, Vietnam.

Nguyen The Vinh, Thai Nguyen University of Information and Communication Technology, Thai Nguyen, Vietnam

Nguyen The Vinh received his PhD in Computer Science from Texas Tech University in 2020. He is currently a lecturer in Software Engineering at the University of Information and Communication Technology- Thai Nguyen University. His main research interests are Computer Science and AI.

Nguyen Huu Cong, Thai Nguyen University, Thai Nguyen, Vietnam

Nguyen Huu Cong received his PhD in automatic control from Hanoi University of Science and Technology in 2003 and was promoted to Associate Professor in 2007. His main research interests are automatic control and optimal control for objects with distributed and slowly changing parameters.

References

Y. Li, G. Hu, J. Du, H. Abbas, and Y. Zhang, “Multitask reading for intelligent legal services,” Future Generation

Computer Systems, vol. 113, pp. 218–227, 2020.

S. Yurii, P. Nataliia, T. Tetiana, P. Tetiana, and H. Stanislav, “Official document as a legal act: Essential aspects,” J. Legal Ethical & Regul. Isses, vol. 22, p. 1, 2019.

F. u. Hassan, T. Le, and X. Lv, “Addressing legal and contractual matters in construction using natural language processing: A critical review,” Journal of Construction Engineering and Management, vol. 147, no. 9, p. 03121004, 2021.

S. R. Ahmad, D. Harris, and I. Sahibzada, “Understanding legal documents: classification of rhetorical role of sentences using deep learning and natural language processing,” in 2020 IEEE 14th International Conference on Semantic Computing (ICSC). IEEE, 2020, pp. 464–467.

G. Marino, D. Licari, P. Bushipaka, G. Comandé, T. Cucinotta et al., “Automatic rhetorical roles classification for

legal documents using legal-transformeroverbert,” in CEUR WORKSHOPPROCEEDINGS,vol.3441. CEUR-WS,2023,

pp. 28–36.

I. Timmer and R. Rietveld, “Rule-based systems for decision support and decision-making in dutch legal practice. a brief overview of applications and implications,” Droit et Societe, vol. 103, pp. 517–534, 11 2019.

D. Bourcier and G. Clergue, “From a rule-based conception to dynamic patterns. analyzing the self-organization of legal systems,” Artificial Intelligence and Law, vol. 7, pp. 211 225, 1999.

K. Pal and J. A. Campbell, “An application of rule-based and case-based reasoning within a single legal knowledge based system,” ACM SIGMIS Database: the DATABASE for Advances in Information Systems, vol. 28, pp. 48–63, 9 1997. [Online]. Available: https://dl.acm.org/doi/10.1145/ 277339.277344

W. Liu, P. Zhou, Z. Zhao, Z. Wang, Q. Ju, H. Deng, and P. Wang, “K-bert: Enabling language representation with

knowledge graph,” Proceedings of the AAAI Conference on Artificial Intelligence, vol. 34, pp. 2901–2908, 4

[Online]. Available: https://ojs.aaai.org/index.php/ AAAI/article/view/5681

C. Xiao, X. Hu, Z. Liu, C. Tu, and M. Sun, “Lawformer: A pre-trained language model for chinese legal long documents,” AI Open, vol. 2, pp. 79–84, 2021. [Online]. Available: https://linkinghub.elsevier.com/retrieve/

pii/S2666651021000176

S. Douka, H. Abdine, M. Vazirgiannis, R. E. Hamdani, and D. R. Amariles, “Juribert: A masked-language model

adaptation for french legal text,” Natural Legal Language Processing Workshop 2021, pp. 95–101, 2021. [Online]. Available: https://aclanthology.org/2021.nllp-1.9.pdf

M. Masala, R. Iacob, A. S. Uban, M.-A. Cidotã, H. Velicu, T. Rebedea, and M. Popescu, “jurbert: A romanian bert model for legal judgement prediction,” Natural Legal Language Processing Workshop 2021, pp. 86–94, 2021. [Online]. Available: https://aclanthology.org/2021.nllp-1.8. pdf

H.-T. Nguyen, M.-K. Phi, X.-B. Ngo, V. Tran, L.-M. Nguyen, and M.-P. Tu, “Attentive deep neural networks

for legal document retrieval,” Artificial Intelligence and Law, vol. 32, pp. 57–86, 3 2024. [Online]. Available:

https://link.springer.com/10.1007/s10506-022-09341-8

H. N. Van, D. Nguyen, P. M. Nguyen, and M. Nguyen, “Miko team: Deep learning approach for legal question answering in alqac 2022,” 2022 14th International Conference on Knowledge and Systems Engineering (KSE), pp. 1–5, 2022. [Online]. Available: https://arxiv.org/pdf/2211.02200

M. Kejriwal, “Ai in industry today,” in Artificial Intelligence for Industries of the Future: Beyond Facebook, Amazon, Microsoft and Google. Springer, 2022, pp. 47–73.

J. Sauvola, S. Tarkoma, M. Klemettinen, J. Riekki, and D. Doermann, “Future of software development with generative ai,” Automated Software Engineering, vol. 31, no. 1, p. 26, 2024.

X. Han, Z. Zhang, N. Ding, Y. Gu, X. Liu, Y. Huo, J. Qiu, Y. Yao, A. Zhang, L. Zhang et al., “Pre-trained models: Past, present and future,” AI Open, vol. 2, pp. 225–250, 2021.

M. Abdin and S. A. Jacobs, “Phi-3 technical report: A highly capable language model locally on your phone,” 4 2024. [Online]. Available: http://arxiv.org/abs/2404.14219

L. Zheng, N. Guha, B. R. Anderson, P. Henderson, and D. E. Ho, “When does pretraining help?: assessing self-supervised learning for law and the casehold dataset of 53,000+ legal holdings,” in Proceedings of the Eighteenth International Conference on Artificial Intelligence and Law. ACM, 6 2021, pp. 159–168.

E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen, “Lora: Low-rank adaptation

of large language models,” 6 2021. [Online]. Available: http://arxiv.org/abs/2106.09685

T. Dettmers, A. Pagnoni, A. Holtzman, and L. Zettlemoyer, “Qlora: Efficient finetuning of quantized llms,” in

Advances in Neural Information Processing Systems, A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt,

and S. Levine, Eds., vol. 36. Curran Associates, Inc., 5 2023, pp. 10088–10115. [Online]. Available:

https://proceedings.neurips.cc/paper_files/paper/2023/file/ 1feb87871436031bdc0f2beaa62a049b-Paper-Conference.pdf

L. Ouyang, J. Wu, and X. Jiang, “Training language models to follow instructions with human feedback,” in Advances in Processing Systems, vol. 35. Neural Information Curran Associates, Inc., 3 2022, pp. 27730–27744. [Online]. Available: https://proceedings.neurips.cc/paper_files/paper/2022/file/ b1efde53be364a73914f58805a001731-Paper-Conference. pdf

I. Chalkidis, A. Jana, D. Hartung, M. Bommarito, I. Androutsopoulos, D. M. Katz, and N. Aletras, “Lexglue:

A benchmark dataset for legal language understanding in english,” Proceedings of the Annual Meeting of

the Association for Computational Linguistics, vol. 1, pp. 4310–4330, 10 2022. [Online]. Available: https:

//arxiv.org/abs/2110.00976v4

Published
2024-11-04