Efficient mining of succinct correlated high utility association rules
Abstract
This paper investigates the problem of mining high utility association rules from quantitative databases. Traditional approaches mainly rely on conventional utility-confidence measures without considering the actual strength of association among items. Due to the absence of a correlation measure, many discovered rules exhibit high utility while the items within the rules co-occur infrequently, making them difficult to interpret and of limited practical significance; moreover, the number of generated rules is often very large. To address these limitations, this paper proposes a novel utility-confidence measure and integrates the Bond correlation measure to ensure that the items in a rule not only yield high utility but also exhibit strong correlations. On this basis, the paper introduces a class of succinct correlated high utility association rules, called SCoHUARs, in which each rule is represented in a compact form that serves as a representative of an equivalence class of rules sharing the same support. In addition, a partitioning method for the set of SCoHUARs is presented to guarantee complete and non-redundant generation of all SCoHUARs, together with several optimization techniques to accelerate the mining process. Finally, the M-SCoHUAR algorithm is designed to efficiently mine SCoHUARs by integrating the proposed theoretical results. Experimental results demonstrate that the proposed algorithm significantly outperforms existing methods in terms of execution time, memory consumption, and the quality of the discovered patterns.
References
R. Agrawal, R. Srikant (1994), “Fast algorithms for mining association rules in large databases”, Proc. of the 20th International Conference on Very Large Data Bases, pp.487- 499.
M. Liu, J. Qu (2012), “Mining high utility itemsets without candidate generation”, Proceedings of ACM International Conference on Information and Knowledge Management, pp.55-64.
B.E. Shie, P.S. Yu, V.S. Tseng (2012), “Mining interesting user behavior patterns in mobile commerce environments”, Applied Intelligence, 38, pp.418-435.
B.E. Shie, C.W. Wu, V.S. Tseng, et al. (2013), “Efficient algorithms for mining high utility itemsets from transactional databases”, EEE Transactions on Knowledge and Data Engineering, 25(8), pp.1772-1786, DOI: 10.1109/ TKDE.2012.59.
Y. Liu, L. Wang, L. Feng (2020), “Mining high utility itemsets based on pattern growth without candidate generation”, Mathmetics, pp.1-22, DOI:10.3390/ math9010035.
M. Liu, J. Qu (2012), “Mining high utility itemsets without candidate generation”, Proceedings of The 21st ACM International Conference on Information and Knowledge Management, pp.55-64, DOI: 10.1145/2396761.2396773.
P.F. Viger, C.W. Wu, S. Zida, et al. (2014), “FHM: Faster high-utility itemset mining using estimated utility cooccurrence pruning”, Lecture Notes in Computer Science, 8502, pp.83-92.
S. Zida, P.F. Viger, J.C.W. Lin, et al. (2015), “A highly efficient algorithm for high-utility itemset mining”, Proceedings of Mexican International Conference on Artificial Intelligence, pp.530-546.
J. Sahoo, A. Kumar, A. Goswami (2015), “An efficient approach for mining association rules from high utility itemsets”, Expert Syst. Appl. 42(13), pp.5754-5778, DOI: 10.1016/j.eswa.2015.02.051.
T. Mai, B. Vo, L.T.T. Nguyen (2017), “A latticebased approach for mining high utility association rules”, Inf. Sci., 399, pp.81-97, DOI: 10.1016/j.ins.2017.02.058.
L.T.T. Nguyen, T. Mai, B. Vo (2019), High Utility Association Rule Mining, High-Utility Pattern Mining, Springer, pp.161-174.
T. Mai, L.T.T. Nguyen, B. Vo, et al. (2020), “Efficient algorithm for mining non-redundant high-utility association rules”, Sensors (Switzerland), 20(4), pp.1-17, DOI: 10.3390/ s20041078.
P.F. Viger, J.C.W. Lin, T. Dinh, et al. (2016), Mining Correlated High-Utility Itemsets Using The Bond Measure, Springer, pp.53-65.
C.C. Aggarwal, P.S. Yu (1998), “New framework for itemset generation, Proc”, ACM Symp. Principles of Database Systems, pp.18-24, DOI: 10.1145/275487.275490.
E.R. Omiecinski (2023), “Alternative interest measures for mining associations in databases”, IEEE Trans. Knowl. Data Eng., pp.57-69.
T. Wu, Y. Chen, J. Han (2010), “Re-examination of interestingness measures in pattern mining: A unified framework”, Data Min. Knowl. Discov., pp.371-397.
T. Tran, H. Duong, T. Truong, et al. (2023), “Efficient mining of concise and informative representations of frequent high utility itemsets”, Eng. Appl. Artif. Intell, 126, DOI: 10.1016/j.engappai.2023.107111.
L.T.T. Nguyen, P. Nguyen, T.D.D. Nguyen, et al. (2019), “Mining high-utility itemsets in dynamic profit databases”, Knowl. Based. Syst., 175, pp.130-144.
T. Truong, H. Duong, B. Le, et al. (2019), “Efficient vertical mining of high average-utility itemsets based on novel upper-bounds”, IEEE Trans. Knowl. Data Eng., 31, pp.301-314.
J.F. Qu, P.F. Viger, M. Liu, et al. (2023), “Mining high utility itemsets using prefix trees and utility vectors”, IEEE Trans. Knowl. Data Eng., 35, pp.10224-10236.
Z. Cheng, W. Fang, W. Shen, et al. (2023), “An efficient utility-list based high-utility itemset mining algorithm”, Applied Intelligence, 53, pp.6992-7006.
R.U. Kiran, P. Veena, P. Ravikumar, et al. (2023), “HDSHUI-miner: a novel algorithm for discovering spatial high-utility itemsets in high-dimensional spatiotemporal databases”, Applied Intelligence, 53, pp.8536-8561.
J.M. Luna, R.U. Kiran, P.F. Viger (2023), “Efficient mining of top-k high utility itemsets through genetic algorithms”, Inf. Sci., 624, pp.529-553.
G.C. Lan, T.P. Hong, V.S. Tseng (2014), “An efficient projection-based indexing approach for mining high utility itemsets”, Knowl. Inf. Syst., 38, pp.85-107.
S. Dawar, V. Goyal, D. Bera (2017), “A hybrid framework for mining high-utility itemsets in a sparse transaction database”, Applied Intelligence, 47, pp.809-827.
D.Q. Huy, P.F. Viger, H. Ramampiaro, et al. (2018), “Efficient high utility itemset mining using buffered utilitylists”, Applied Intelligence, 48, pp.1859-1877.
Z. Deng (2018), “An efficient structure for fast mining high utility itemsets”, Applied Intelligence, 48, pp.3161-3177.
T. Truong, H. Duong, B. Le, et al. (2019), “Efficient high average-utility itemset mining using novel vertical weak upper-bounds”, Knowl. Based. Syst., 183, DOI: 1016/j. knosys.2019.07.018.
H. Duong, T. Hoang, T. Tran, et al. (2022), “Efficient algorithms for mining closed and maximal high utility itemsets”, Knowl. Based. Syst., 257, DOI: 10.1016/j. knosys.2022.109921.
V.S. Tseng, C. Wu, P.F. Viger (2015), “Efficient algorithms for mining the concise and lossless representation of high utility itemsets”, IEEE Trans. Knowl. Data Eng., 27, pp.726-739.
C.W. Wu, P.F. Viger, J.Y. Gu, et al. (2015), “Mining closed + high utility itemsets without candidate generation”, Conference on Technologies and Applications of Artificial Intelligence, pp.187-194.
P.F. Viger, S. Zida, J.C.W. Lin, et al. (2016), “EFIMClosed: Fast and memory efficient discovery of closed high-utility itemsets”, International Conference on Machine Learning and Data Mining in Pattern Recognition, pp.199- 213.
L.T.T. Nguyen, V.V. Vu, M.T.H. Lam, et al. (2019), “An efficient method for mining high utility closed itemsets”, Inf. Sci., 495, pp.78-99.
P.F. Viger, C.W. Wu, V.S. Tseng (2014), Novel Concise Representations of High Utility Itemsets Using Generator Patterns, International Conference on Advanced Data Mining and Applications, pp.30-43.
J. Sahoo, A.K. Das, A. Goswami (2014), “An algorithm for mining high utility closed itemsets and generators”, Proceedings of The 2014 IEEE International Conference on Data Science and Advanced Analytics, pp.1-9.
S. Bouasker, S. Yahia (2015), “Key correlation mining by simultaneous monotone and anti-monotone constraints checking”, Proceedings of The ACM Symposium on Applied Computing, pp.851-856.
P.F. Viger, A. Gomariz, A. Soltani, et al. (2014), “SPMF: A java open-source pattern mining library”, J. Mach. Learn. Res., 15, pp.3569-3573.
