Thuật toán song song khai thác itemset lợi nhuận phổ biến Skyline
Parallel Algorithm Exploits Skyline Common Interest Element Set
Abstract
Các itemset lợi nhuận phổ biến Skyline (SFUI) có thể cung cấp nhiều thông tin hữu ích hơn cho việc ra quyết định với việc xem xét cả hai yếu tố là tần suất xuất hiện và lợi nhuận của chúng. Kể từ khi bài toán khai thác itemset lợi nhuận phổ biến Skyline được Goyal V. và các cộng sự đề xuất vào năm 2015, đến nay đã có nhiều thuật toán tuần tự được đề xuất nhằm cải thiện hiệu suất khai thác, tuy nhiên hầu hết các thuật toán đều có hiệu suất kém khi khai thác các tập dữ liệu lớn phổ biến hiện nay. Trong bài báo này, chúng tôi đề xuất một thuật toán song song có tên là ParaSFUI-UF dựa trên thuật toán tuần tự SFUI-UF là một thuật toán khai thác itemset lợi nhuận phổ biến Skyline hiệu quả nhất hiện nay. Các kết quả thực nghiệm cho thấy thuật toán ParaSFUI-UF vượt trội so với thuật toán SFUI-UF.
References
Vikram Goyal, Ashish Sureka, Dhaval Patel, "Efficient skyline itemsets mining", C3S2E ’15: Proceedings of the Eighth International C* Conference on Computer Science & Software Engineering (2015), 119–124.
Pan Jeng-Shyanga, Lin Jerry Chun-Weib, Yang Lub, FournierViger Philippec, Hong Tzung-Peid, "Efficiently mining of skyline frequent-utility patterns", Intelligent Data Analysis, vol. 21, no. 6 (2017), 1407-1423.
Jerry Chun-Wei Lin, Lu Yang, Philippe Fournier-Viger, Tzung-Pei Hong, "Mining of skyline patterns by considering both frequent and utility constraints", Engineering Applications of Artificial Intelligence, Volume 77 (2019), 229-238.
Hung Manh Nguyen; Anh Viet Phan; Lai Van Pham: FSKYMINE, "A Faster Algorithm For Mining Skyline Frequent Utility Itemsets", Proceedings of the 6th NAFOSTED Conference on Information and Computer Science, (2019), 251–255.
Wei Song, Chuanlong Zheng, Philippe Fournier-Viger, "Mining Skyline Frequent-Utility Itemsets with Utility Filtering", 18th Pacific Rim International Conference on Artificial Intelligence, PRICAI (2021), Part I ,411–424.
Bertil Schmidt, Jorge Gonzalez-Martinez, Christian Hundt, Moritz Schlarb, Parallel Programming: Concepts and Practice, Morgan Kaufmann (2017).
Julian Shun: Shared-Memory Parallelism Can Be Simple, Fast, and Scalable, Association for Computing Machinery and Morgan & Claypool (2017).
Ying Liu, Wei-keng Liao, Alok Choudhary, "A two-phase algorithm for fast discovery of high utility itemsets", Advances in Knowledge Discovery and Data Mining, 9th Pacific-Asia Conference, PAKDD (2005), 689–695.
Mengchi Liu, Junfeng Qu: "Mining high utility itemsets without candidate generation", 21st ACM International Conference on Information and Knowledge Management, (2012), 55–64.
Jerry Chun-Wei Lin, Lu Yang, Philippe Fournier-Viger, Siddharth Dawar, Vikram Goyal, Ashish Sureka, and Bay Vo, "A more efficient algorithm to mine skyline frequentutility patterns", In International Conference on Genetic and Evolutionary Computing, 2017, 127–135.
Bay Vo, Loan T. T. Nguyen, Trinh D. D. Nguyen, Philippe Fournier-Viger, Unil Yun, "A Multi-Core Approach to Efficiently Mining High-Utility Itemsets in Dynamic Profit Databases", IEEE Access (2020), 85890 – 85899.
Mohammed J. Zaki, "Parallel Sequence Mining on Shared-Memory Machines", Journal of Parallel and Distributed Computing, Volume 61, Issue 3 (2001), 401-426.
Krishan Kumar Sethi, Dharavath Ramesh , Damodar Reddy Edla: P-FHM+, "Parallel high utility itemset mining algorithm for big data processing", Procedia Computer Science, Volume 132 (2018), 918-927.
O. Jamsheela, G. Raju, "Parallelization of Frequent Itemset Mining Methods with FP-tree: An Experiment with PrePost+ Algorithm", The International Arab Journal of Information Technology, Vol. 18, No. 2(2021), 208-231.
An Open-Source Data Mining Library: http://www.philippefournier-viger.com/spmf/
