SiMoI New Method to Solve the Sparsity Problem in Collaborative Filtering
Abstract
Sparsity data is a major challenge in collaborative recommendation systems, characterized by the predominance of missing values within the user-item matrix. When a substantial portion of data is unavailable, the estimation process becomes hindered, and prediction accuracy declines due to limited usable information. To address this issue, this study introduces a novel method called SiMoI (Similarity, Mode, and Minimum Imputation), which is adaptively designed to handle high levels of sparsity. The SiMoI method combines user similarity with imputation strategies based on mode and minimum values. By leveraging subsets of the most informative users and items, the method efficiently fills missing entries while maintaining prediction stability. Evaluation was conducted using both real and synthetic datasets with varying sizes and degrees of sparsity, including an extreme scenario with 93.7% missing data. Experimental results show that SiMoI consistently produces more accurate predictions than baseline methods. Under high-sparsity conditions, SiMoI achieved an RMSE as low as 0.823, outperforming KNNI (0.947) and MEAN (1.021). Moreover, SiMoI demonstrated resilience across different data scales and sparsity distributions, indicating its flexibility and scalability in diverse contexts. These findings suggest that SiMoI is an effective and stable approach for addressing sparsity and holds strong potential for implementation in user-based recommendation systems, particularly in real-world scenarios where data availability is frequently limited.
Keywords
Full Text:
PDFReferences
J. Lu, D. Wu, M. Mao, W. Wang, and G. Zhang, “Recommender system application developments: A survey,” Decis Support Syst, vol. 74, pp. 12–32, Jun. 2015, doi: 10.1016/j.dss.2015.03.008.
B. Patel, P. Desai, and U. Panchal, “Methods of Recommender System: A Review,” in International Conference on Innovations in information Embedded and Communication Systems (ICIIECS), 2017, pp. 1–4.
L. Yao, Q. Z. Sheng, A. H. H. Ngu, J. Yu, and A. Segev, “Unified collaborative and content-based web service recommendation,” IEEE Trans Serv Comput, vol. 8, no. 3, pp. 453–466, May 2015, doi: 10.1109/TSC.2014.2355842.
Y. Xu and J. Yin, “Collaborative recommendation with user generated content,” Eng Appl Artif Intell, vol. 45, pp. 281–294, Oct. 2015, doi: 10.1016/j.engappai.2015.07.012.
P. B. Thorat, R. M. Goudar, and S. Barve, “Survey on Collaborative Filtering, Content-based Filtering and Hybrid Recommendation System,” Int J Comput Appl, vol. 110, no. 4, pp. 975–8887, 2015.
D. M. J. S. Bowman, C. Haverkamp, K. D. Rann, and L. D. Prior, “Differential demographic filtering by surface fires: How fuel type and fuel load affect sapling mortality of an obligate seeder savanna tree,” Journal of Ecology, vol. 106, no. 3, pp. 1010–1022, May 2018, doi: 10.1111/1365-2745.12819.
A. Fareed, S. Hassan, S. B. Belhaouari, and Z. Halim, “A collaborative filtering recommendation framework utilizing social networks,” Machine Learning with Applications, vol. 14, p. 100495, Dec. 2023, doi: 10.1016/j.mlwa.2023.100495.
G. Behera and N. Nain, “Collaborative Filtering with Temporal Features for Movie Recommendation System,” in Procedia Computer Science, Elsevier B.V., 2022, pp. 1366–1373. doi: 10.1016/j.procs.2023.01.115.
Z. Ren, B. Peng, T. K. Schleyer, and X. Ning, “Hybrid collaborative filtering methods for recommending search terms to clinicians,” J Biomed Inform, vol. 113, Jan. 2021, doi: 10.1016/j.jbi.2020.103635.
K. Kobyshev, N. Voinov, and I. Nikiforov, “Hybrid image recommendation algorithm combining content and collaborative filtering approaches,” in Procedia Computer Science, Elsevier B.V., 2021, pp. 200–209. doi: 10.1016/j.procs.2021.10.020.
Y. Afoudi, M. Lazaar, and M. Al Achhab, “Hybrid recommendation system combined content-based filtering and collaborative prediction using artificial neural network,” Simul Model Pract Theory, vol. 113, Dec. 2021, doi: 10.1016/j.simpat.2021.102375.
F. E. Alsaadi, Z. Wang, N. S. Alharbi, Y. Liu, and N. D. Alotaibi, “A new framework for collaborative filtering with p-moment-based similarity measure: Algorithm, optimization and application,” Knowl Based Syst, vol. 248, Jul. 2022, doi: 10.1016/j.knosys.2022.108874.
K. Vahidy Rodpysh, S. J. Mirabedini, and T. Banirostam, “Resolving cold start and sparse data challenge in recommender systems using multi-level singular value decomposition,” Computers and Electrical Engineering, vol. 94, Sep. 2021, doi: 10.1016/j.compeleceng.2021.107361.
S. Ahmadian, N. Joorabloo, M. Jalili, and M. Ahmadian, “Alleviating data sparsity problem in time-aware recommender systems using a reliable rating profile enrichment approach,” Expert Syst Appl, vol. 187, Jan. 2022, doi: 10.1016/j.eswa.2021.115849.
Mohamad Fahmi Hafidz and Sri Lestari, “Solution to Scalability and Sparsity Problems in Collaborative Filtering using K-Means Clustering and Weight Point Rank (WP-Rank),” Jurnal RESTI (Rekayasa Sistem dan Teknologi Informasi), vol. 7, no. 4, pp. 743–750, Aug. 2023, doi: 10.29207/resti.v7i4.4543.
M. Singh, “Scalability and sparsity issues in recommender datasets: a survey,” Knowl Inf Syst, vol. 62, no. 1, pp. 1–43, Jan. 2020, doi: 10.1007/s10115-018-1254-2.
F. O. Isinkaye, Y. O. Folajimi, and B. A. Ojokoh, “Recommendation systems: Principles, methods and evaluation,” Nov. 01, 2015, Elsevier B.V. doi: 10.1016/j.eij.2015.06.005.
G. A. J. Satvika, I. N. Sukajaya, and I. G. A. Gunadi, “Improving k-nearest neighbor performance using permutation feature importance to predict student success in study,” Indonesian Journal of Electrical Engineering and Computer Science, vol. 35, no. 3, pp. 1835–1844, Sep. 2024, doi: 10.11591/ijeecs.v35.i3.pp1835-1844.
A. Fadlil, Herman, and D. Praseptian M, “K Nearest Neighbor Imputation Performance on Missing Value Data Graduate User Satisfaction,” Jurnal RESTI (Rekayasa Sistem dan Teknologi Informasi), vol. 6, no. 4, pp. 570–576, Aug. 2022, doi: 10.29207/resti.v6i4.4173.
A. Fadlil, Herman, and M. Dikky Praseptian, “Single Imputation Using Statistics-Based and K Nearest Neighbor Methods for Numerical Datasets,” Ingenierie des Systemes d’Information, vol. 28, no. 2, pp. 451–459, Apr. 2023, doi: 10.18280/isi.280221.
H. Shi, P. Wang, X. Yang, and H. Yu, “An Improved Mean Imputation Clustering Algorithm for Incomplete Data,” Neural Process Lett, vol. 54, no. 5, pp. 3537–3550, Oct. 2022, doi: 10.1007/s11063-020-10298-5.
M. Wolbers, A. Noci, P. Delmar, C. Gower-Page, S. Yiu, and J. W. Bartlett, “Standard and reference-based conditional mean imputation,” Pharm Stat, vol. 21, no. 6, pp. 1246–1257, Nov. 2022, doi: 10.1002/pst.2234.
G. Jain, T. Mahara, A. Kumar, and S. C. Sharma, “Time-Aware Based Recommendation System using Gower’s Coefficients: Enhancing Personalized Recommendation,” in Procedia Computer Science, Elsevier B.V., 2024, pp. 3379–3388. doi: 10.1016/j.procs.2024.04.318.
A. Jadhav, D. Pramod, and K. Ramanathan, “Comparison of Performance of Data Imputation Methods for Numeric Dataset,” Applied Artificial Intelligence, vol. 33, no. 10, pp. 913–933, Aug. 2019, doi: 10.1080/08839514.2019.1637138.
DOI: https://doi.org/10.47738/jads.v7i1.1015
Refbacks
- There are currently no refbacks.

Journal of Applied Data Sciences
| ISSN | : | 2723-6471 (Online) |
| Publisher | : | Bright Publisher |
| Website | : | http://bright-journal.org/JADS |
| : | taqwa@amikompurwokerto.ac.id (principal contact) | |
| support@bright-journal.org (technical issues) |
This work is licensed under a Creative Commons Attribution-ShareAlike 4.0




.png)