Khai thác hiệu quả các luật kết hợp súc tích có lợi ích và tương quan cao

  • Trần Thống
  • Hoàng Minh Tiến
  • Dương Văn Hải
  • Trương Chí Tín
##plugins.pubIds.doi.readerDisplayName##: https://doi.org/10.31276/CNTT.2026.1433
Keywords: độ đo tương quan Bond; hàm tin cậy - lợi ích; luật kết hợp lợi ích cao; luật kết hợp súc tích; luật kết hợp tương quan cao; mẫu đóng; mẫu sinh

Abstract

Bài báo này xét bài toán khai thác các luật kết hợp lợi ích cao từ cơ sở dữ liệu định lượng. Các tiếp cận truyền thống chủ yếu dựa trên độ tin cậy - lợi ích thông thường mà không xem xét mức độ gắn kết thực sự giữa các mục. Do thiếu một độ đo tương quan, nhiều luật được phát hiện có lợi ích cao nhưng các mục trong luật lại ít đồng xuất hiện, khó diễn giải và kém ý nghĩa thực tiễn; đồng thời số lượng luật sinh ra thường rất lớn. Để khắc phục các hạn chế này, bài báo đề xuất một độ đo tin cậy dựa trên lợi ích mới, đồng thời tích hợp thêm độ đo tương quan Bond nhằm bảo đảm các mục trong luật vừa mang lại lợi ích cao vừa có mối liên kết chặt chẽ. Trên nền tảng đó, bài báo đề xuất lớp luật kết hợp súc tích có lợi ích cao và tương quan mạnh, gọi là SCoHUAR, trong đó mỗi luật được biểu diễn dưới dạng rút gọn, đại diện cho một nhóm luật tương đương về độ hỗ trợ. Ngoài ra, bài báo còn trình bày phương pháp phân hoạch tập các SCoHUAR giúp bảo đảm sinh đầy đủ và không trùng lặp các luật SCoHUAR và các kỹ thuật tối ưu để tăng tốc quá trình khai thác. Cuối cùng, thuật toán M-SCoHUAR được thiết kế để khai thác hiệu quả các SCoHUAR, trong đó tích hợp các kết quả lý thuyết được đề xuất. Các kết quả thực nghiệm chỉ ra rằng thuật toán được đề xuất có nhiều ưu điểm nổi bật hơn so với các phương pháp trước đây về thời gian thực thi, mức bộ nhớ sử dụng và chất lượng của các mẫu được khám phá.

Author Biographies

Trần Thống

Khoa Công nghệ thông tin, Trường Đại học Khoa học Tự nhiên, TP. Hồ Chí Minh, 227 đường Nguyễn Văn Cừ, phường Chợ Quán, TP. Hồ Chí Minh, Việt Nam

Đại học Quốc gia TP. Hồ Chí Minh, phường Linh Xuân, TP. Hồ Chí Minh, Việt Nam

Khoa Công nghệ thông tin, Trường Đại học Đà Lạt, Đường Phù Đổng Thiên Vương, Phường Lâm Viên, TP. Đà Lạt, tỉnh Lâm Đồng, Việt Nam

Hoàng Minh Tiến

Khoa Công nghệ thông tin, Trường Đại học Khoa học Tự nhiên, TP. Hồ Chí Minh, 227 đường Nguyễn Văn Cừ, phường Chợ Quán, TP. Hồ Chí Minh, Việt Nam

Đại học Quốc gia TP. Hồ Chí Minh, phường Linh Xuân, TP. Hồ Chí Minh, Việt Nam

Khoa Toán - Tin học, Trường Đại học Đà Lạt, Đường Phù Đổng Thiên Vương, Phường Lâm Viên, TP. Đà Lạt, tỉnh Lâm Đồng, Việt Nam

Dương Văn Hải

Khoa Toán - Tin học, Trường Đại học Đà Lạt, Đường Phù Đổng Thiên Vương, Phường Lâm Viên, TP. Đà Lạt, tỉnh Lâm Đồng, Việt Nam

Trương Chí Tín

Khoa Toán - Tin học, Trường Đại học Đà Lạt, Đường Phù Đổng Thiên Vương, Phường Lâm Viên, TP. Đà Lạt, tỉnh Lâm Đồng, Việt Nam

References

R. Agrawal, R. Srikant (1994), “Fast algorithms for mining association rules in large databases”, Proc. of the 20th International Conference on Very Large Data Bases, pp.487- 499.

M. Liu, J. Qu (2012), “Mining high utility itemsets without candidate generation”, Proceedings of ACM International Conference on Information and Knowledge Management, pp.55-64.

B.E. Shie, P.S. Yu, V.S. Tseng (2012), “Mining interesting user behavior patterns in mobile commerce environments”, Applied Intelligence, 38, pp.418-435.

B.E. Shie, C.W. Wu, V.S. Tseng, et al. (2013), “Efficient algorithms for mining high utility itemsets from transactional databases”, EEE Transactions on Knowledge and Data Engineering, 25(8), pp.1772-1786, DOI: 10.1109/ TKDE.2012.59.

Y. Liu, L. Wang, L. Feng (2020), “Mining high utility itemsets based on pattern growth without candidate generation”, Mathmetics, pp.1-22, DOI:10.3390/ math9010035.

M. Liu, J. Qu (2012), “Mining high utility itemsets without candidate generation”, Proceedings of The 21st ACM International Conference on Information and Knowledge Management, pp.55-64, DOI: 10.1145/2396761.2396773.

P.F. Viger, C.W. Wu, S. Zida, et al. (2014), “FHM: Faster high-utility itemset mining using estimated utility cooccurrence pruning”, Lecture Notes in Computer Science, 8502, pp.83-92.

S. Zida, P.F. Viger, J.C.W. Lin, et al. (2015), “A highly efficient algorithm for high-utility itemset mining”, Proceedings of Mexican International Conference on Artificial Intelligence, pp.530-546.

J. Sahoo, A. Kumar, A. Goswami (2015), “An efficient approach for mining association rules from high utility itemsets”, Expert Syst. Appl. 42(13), pp.5754-5778, DOI: 10.1016/j.eswa.2015.02.051.

T. Mai, B. Vo, L.T.T. Nguyen (2017), “A latticebased approach for mining high utility association rules”, Inf. Sci., 399, pp.81-97, DOI: 10.1016/j.ins.2017.02.058.

L.T.T. Nguyen, T. Mai, B. Vo (2019), High Utility Association Rule Mining, High-Utility Pattern Mining, Springer, pp.161-174.

T. Mai, L.T.T. Nguyen, B. Vo, et al. (2020), “Efficient algorithm for mining non-redundant high-utility association rules”, Sensors (Switzerland), 20(4), pp.1-17, DOI: 10.3390/ s20041078.

P.F. Viger, J.C.W. Lin, T. Dinh, et al. (2016), Mining Correlated High-Utility Itemsets Using The Bond Measure, Springer, pp.53-65.

C.C. Aggarwal, P.S. Yu (1998), “New framework for itemset generation, Proc”, ACM Symp. Principles of Database Systems, pp.18-24, DOI: 10.1145/275487.275490.

E.R. Omiecinski (2023), “Alternative interest measures for mining associations in databases”, IEEE Trans. Knowl. Data Eng., pp.57-69.

T. Wu, Y. Chen, J. Han (2010), “Re-examination of interestingness measures in pattern mining: A unified framework”, Data Min. Knowl. Discov., pp.371-397.

T. Tran, H. Duong, T. Truong, et al. (2023), “Efficient mining of concise and informative representations of frequent high utility itemsets”, Eng. Appl. Artif. Intell, 126, DOI: 10.1016/j.engappai.2023.107111.

L.T.T. Nguyen, P. Nguyen, T.D.D. Nguyen, et al. (2019), “Mining high-utility itemsets in dynamic profit databases”, Knowl. Based. Syst., 175, pp.130-144.

T. Truong, H. Duong, B. Le, et al. (2019), “Efficient vertical mining of high average-utility itemsets based on novel upper-bounds”, IEEE Trans. Knowl. Data Eng., 31, pp.301-314.

J.F. Qu, P.F. Viger, M. Liu, et al. (2023), “Mining high utility itemsets using prefix trees and utility vectors”, IEEE Trans. Knowl. Data Eng., 35, pp.10224-10236.

Z. Cheng, W. Fang, W. Shen, et al. (2023), “An efficient utility-list based high-utility itemset mining algorithm”, Applied Intelligence, 53, pp.6992-7006.

R.U. Kiran, P. Veena, P. Ravikumar, et al. (2023), “HDSHUI-miner: a novel algorithm for discovering spatial high-utility itemsets in high-dimensional spatiotemporal databases”, Applied Intelligence, 53, pp.8536-8561.

J.M. Luna, R.U. Kiran, P.F. Viger (2023), “Efficient mining of top-k high utility itemsets through genetic algorithms”, Inf. Sci., 624, pp.529-553.

G.C. Lan, T.P. Hong, V.S. Tseng (2014), “An efficient projection-based indexing approach for mining high utility itemsets”, Knowl. Inf. Syst., 38, pp.85-107.

S. Dawar, V. Goyal, D. Bera (2017), “A hybrid framework for mining high-utility itemsets in a sparse transaction database”, Applied Intelligence, 47, pp.809-827.

D.Q. Huy, P.F. Viger, H. Ramampiaro, et al. (2018), “Efficient high utility itemset mining using buffered utilitylists”, Applied Intelligence, 48, pp.1859-1877.

Z. Deng (2018), “An efficient structure for fast mining high utility itemsets”, Applied Intelligence, 48, pp.3161-3177.

T. Truong, H. Duong, B. Le, et al. (2019), “Efficient high average-utility itemset mining using novel vertical weak upper-bounds”, Knowl. Based. Syst., 183, DOI: 1016/j. knosys.2019.07.018.

H. Duong, T. Hoang, T. Tran, et al. (2022), “Efficient algorithms for mining closed and maximal high utility itemsets”, Knowl. Based. Syst., 257, DOI: 10.1016/j. knosys.2022.109921.

V.S. Tseng, C. Wu, P.F. Viger (2015), “Efficient algorithms for mining the concise and lossless representation of high utility itemsets”, IEEE Trans. Knowl. Data Eng., 27, pp.726-739.

C.W. Wu, P.F. Viger, J.Y. Gu, et al. (2015), “Mining closed + high utility itemsets without candidate generation”, Conference on Technologies and Applications of Artificial Intelligence, pp.187-194.

P.F. Viger, S. Zida, J.C.W. Lin, et al. (2016), “EFIMClosed: Fast and memory efficient discovery of closed high-utility itemsets”, International Conference on Machine Learning and Data Mining in Pattern Recognition, pp.199- 213.

L.T.T. Nguyen, V.V. Vu, M.T.H. Lam, et al. (2019), “An efficient method for mining high utility closed itemsets”, Inf. Sci., 495, pp.78-99.

P.F. Viger, C.W. Wu, V.S. Tseng (2014), Novel Concise Representations of High Utility Itemsets Using Generator Patterns, International Conference on Advanced Data Mining and Applications, pp.30-43.

J. Sahoo, A.K. Das, A. Goswami (2014), “An algorithm for mining high utility closed itemsets and generators”, Proceedings of The 2014 IEEE International Conference on Data Science and Advanced Analytics, pp.1-9.

S. Bouasker, S. Yahia (2015), “Key correlation mining by simultaneous monotone and anti-monotone constraints checking”, Proceedings of The ACM Symposium on Applied Computing, pp.851-856.

P.F. Viger, A. Gomariz, A. Soltani, et al. (2014), “SPMF: A java open-source pattern mining library”, J. Mach. Learn. Res., 15, pp.3569-3573.

Published
2025-12-20