Impact of Sample Size on the Robustness of Machine Learning Algorithms for Detecting Loan Defaults Using Imbalanced Data
Abstract
This study aimed to assess the impact of sample size on the robustness of five machine learning classifiers: Support Vector Machine (SVM), Random Forest (RF), Naïve Bayes (NB), Decision Trees (DT), and K-Nearest Neighbour (K-NN). Although there are data-balancing techniques that aid in addressing data imbalance, they have some limitations which are discussed in this paper. The current study continues the trend in the application of these five ML classifiers for credit default detection, but it makes a contribution by examining whether sample size increment can better their performance when they are trained using a different imbalanced loan default dataset which has not been the focus of previous studies, although most ML algorithms are known to perform well when trained with large datasets. The study used a secondary loan default imbalanced dataset from Kaggle.com, where 85% of participants made loan payments and 15% defaulted. Stratified random sampling was used to select different sample sizes starting with 2% of the total observations, followed by 5%, then 10% up to 90% of the dataset, with the dependent variable being the stratum. The study found no consistent change in the classification metrics with the change in sample size, but RF and DT achieved 100% performance regardless of sample size and are therefore recommended as the most robust to data imbalance in loan default detection. The average classification metrics for NB and K-NN ranged from 72% to 92%, and SVM produced the lowest averages which were between 69% and 75%. NB, K-NN and SVM yielded poor sensitivity rates of 0% to 53%, indicating poor loan payments prediction but they had sensitivity scores in range of 84% to 86%, indicating good loan default classification. Future studies should consider other sampling methods, deep and hybrid learning methods with comparison to RF and DT.
Keywords
Full Text:
PDFReferences
Ashraf, A.S. and T. Ahmed, "Machine learning shrewd approach for an imbalanced dataset conversion samples
", Journal of Engineering and Technology, 11, 2020.
López, V., A. Fernández, S. García, V. Palade, and F. Herrera, "An Insight into Classification with Imbalanced Data: Empirical Results and Current Trends on Using Data Intrinsic Characteristics. ", Inf. Sci, 113−141, 2013.
Zheng, W. and M. Jin, "The Effects of Class Imbalance and Training Data Size on Classifier Learning: An Empirical Study", SN Computer Science 1, 2020.
Dube, L. and T. Verster, "Enhancing classification performance in imbalanced datasets: A comparative analysis of machine learning models", Data Science in Finance and Economics, 3, 2023.
Vuttipittayamongkol, P., E. Elyan, and A. Petrovski, "On the class overlap problem in imbalanced data classification", Knowledge-based systems [online], 212, 2021.
Loo, W.T., K.W. Khaw, X.Y. Chew, A. Alnoor, and S.T. Lim, "Predicting loan default using machine learning algorithms: A case study in India", Journal of Engineering and Technology, 14, 17-27, 2023.
Elrahman, S.M.A. and A. Abraham, "A Review of Class Imbalance Problem", Journal of Network and Innovative Computing, 1, 332-340, 2013.
Singh, A., R.K. Ranjan, and A. Tiwari, "Credit Card Fraud Detection under Extreme Imbalanced Data: A Comparative Study of Data-level Algorithms", Journal of Experimental & Theoretical Artificial Intelligence, 34, 571-598, 2022.
Madaan, M., A. Kumar, C. Keshri, R. Jain, and P. Nagrath, "Loan default prediction using decision trees and random forest: A comparative study", IOP Conference Series: Materials Science and Engineering, 1-12, 2021.
Madaan, M., A. Kumar, C. Keshri, R. Jain, and P. Nagrath, "Loan default prediction using decision trees and random forest:A comparative study ", ICCRDA 2020, 1-12, 2021.
Anand, M., A. Velu, and P. Whig, "Prediction of Loan Behaviour with Machine Learning Models for Secure Banking", Journal of Computer science an Engineering (JCSE), 3, 1-13, 2022.
Shokeen, D., V. Grover, and V. Verma, "Mitigating loan default risk in the banking sector: Machine learning solutions and comparative performance analysis", Research gate, 2023.
Fati, S.M., "A loan default prediction model using machine learning and feature engineering", ICIC Express letters, 18, 27-37, 2024.
Andrić, K., D. Kalpić, and Z. Bohaček, "An insight into the effects of class imbalance and sampling on classification accuracy in credit risk assessment", Computer Science and Information Systems, 16, 155-178, 2019.
Ramezan, C.A., T.A. Warner, A.E. Maxwell, and B.S. Price, "Effects of Training Set Size on Supervised Machine-Learning Land-Cover Classification of Large-Area High-Resolution Remotely Sensed Data", Electronics MDPI, 13, 2021.
Peling, I.B.A., N. Arnawan, P.A. Arthawan, and I. Janardana, "Implementation of Data Mining To Predict Period of Students Study Using Naive Bayes Algorithm", International Journal of Engineering and Emerging Technology, 2, 53-59, 2017.
Batta, M., "Machine Learning Algorithms - A Review", International Journal of Science and Research (IJSR), 9, 381-386, 2018.
Heydaria, S.S. and G. Mountrakisa, "Effect of classifier selection,reference sample size,reference class distribution and scene heterogeneity in per-pixel classification accuracy using 26 Landsat sites", Elsevier, 2017.
Punia, S.K., T. Stephan, and R. Patan, "Performance Analysis of Machine Learning Algorithms for Big Data Classification:ML and AI-Based Algorithms for Big Data Analysis", International Journal of E-Health and Medical Communications, 12, 2021.
Chen, Z., C. Li, and W. Sun, "Bitcoin price prediction using machine learning: An approach
to sample dimension engineering", Journal of Computational and Applied Mathematics, 2020.
Hartmann, J., J. Huppertz, C. Schamp, and M. Heitmann, "Comparing automated text classification methods", International Journal of Research in Marketing, 36, 20-38, 2019.
Karatas, G., O. Demir, and O.K. Sahingoz, "Increasing the Performance of Machine Learning-Based IDSs on an Imbalanced and Up-to-Date Dataset", IEEE Access, 8, 2020.
Li, Q., C. Zhao, X. He, K. Chen, and R. Wang, "The Impact of Partial Balance of Imbalanced Dataset on Classification Performance", Electronics MDPI, 11, 2022.
Guo, Q., J. Zhang, S. Guo, Z. Ye, H. Deng, X. Hou, and H. Zhang, "Urban Tree Classification Based on Object-Oriented Approach and Random Forest Algorithm Using Unmanned Aerial Vehicle (UAV) Multispectral Imagery", Remote Sens, 2022.
Parsa, A.B., H. Taghipour, S. Derrible, and A.K. Mohammadian, "Real-time accident detection: coping with imbalanced data.", Elsevier, 2019.
Sheykhmousa, M., M. Mahdianpari, H. Ghanbari, F. Mohammadimanesh, P. Ghamisi, and S. Homayouni, "Support Vector Machine Versus Random Forest for Remote Sensing Image Classification: A Meta-Analysis and Systematic Review", IEEE JOURNAL OF SELECTED TOPICS IN APPLIED EARTH OBSERVATIONS AND REMOTE SENSING, , 13, 2020.
Pinheiro, A.A., L.M. Brandao, and C. Da Costa, "Vibration Analysis of Rotary Machines Using Machine Learning Techniques", EJERS, European Journal of Engineering Research and Science, 4, 2019.
Fuqing, Y., U. Kumar, and D. Galar, "A comparative study of artificial neural networks and support vector machine for fault diagnosis ", International Journal of Performability Engineering, 9, 49-60, 2013.
Flores, V. and C. Leiva, "A Comparative Study on Supervised Machine Learning Algorithms for Copper Recovery Quality Prediction in a Leaching Process", Sensors
, 2021.
Saljoughi, B.S. and A. Hezarkhani, "A comparative analysis of artificial neural network (ANN), wavelet neural network (WNN), and support vector machine (SVM) data–driven models to mineral potential mapping for copper mineralization in the Shahr-e-Babak region, Kerman, Iran. ", Appl. Geomat., 10, 229–256., 2018.
Zhang , F., M. Melissa Petersen, L. Leigh Johnson, J. Hall, and S.E. O’bryant, "Hyperparameter Tuning with High Performance Computing Machine Learning for Imbalanced Alzheimer’s Disease Data", Appl. Sci, 12, 2022.
Wainer, J. and P. Fonseca, "How to tune the RBF SVM hyperparameters?:An empirical evaluation of 18 search algorithms", 1, 2020.
Eesa, A.S., Z. Orman, and A.M.A. Brifcani, "A novel feature-selection approach based on the cuttlefish optimization algorithm for intrusion detection systems", Expert Systems with Applications, 42, 2670–2679, 2015.
Liang, J., Z. Qin, S. Xiao, L. Ou, and X. Lin, "Efficient and secure decision tree classification for cloud-assisted online diagnosis services", IEEE Transactions on Dependable and Secure Computing, 2019.
Yang, F., "An Extended Idea about Decision Trees", International Conference on Computational Science and Computational Intelligence (CSCI), 349–354, 2019.
Nai-Arun, N. and R. Moungmai, "Comparison of Classifiers for the Risk of Diabetes Prediction", Elsevier, 2015.
Dey, A., "Machine learning algorithms: a review", International Journal of Computer Science and Information Technologies, 7, 1174–1179, 2016.
Jijo, B.T. and A.M. Abdulazeez, "Classification Based on Decision Tree Algorithm for Machine Learning", Journal of Applied Science and Technology Trends, 2, 20 – 28, 2021.
Awoyemi, J.O., A.O. Adetunmbi, and S.A. Oluwadare, "Credit card fraud detection using machine learning techniques: A comparative analysis", 2017 International Conference on Computing Networking and Informatics (ICCNI), Lagos, 1-9, 2017.
Safitri, A.R. and M.A. Muslim, "Improved Accuracy of Naive Bayes Classifier for Determination of Customer Churn Uses SMOTE and Genetic Algorithms", Journal of soft computing and exploration, 6 2020.
Sateesh, N., E, S. Bhanusri, K, K.M. Pasha, S.K. Sameer, P.G. Krishna, and Sunitha.A, "Crop Recommendation system using machine learning algorithm ", UGC Care Group I, 13, 187, 2023.
Rahman, A.K.M., F.M. Javed Mehedi Shamrat, Z. Tasnim, J. Roy, and S.A. Hossain, "A Comparative Study On Liver Disease Prediction Using Supervised Machine Learning Algorithms", International journal of scientific & technology research 8, 2019.
Wang, Q., S. Wang, B. Wei, W. Chen, and Y. Zhang, "Weighted K-NN Classification Method of Bearings Fault Diagnosis With Multi-Dimensional Sensitive Features", IEEE, 9, 2021.
Foysal, K.H., H.J. Chang, F. Bruess, and J.W. Chong, "SmartFit: Smartphone Application for Garment Fit Detection", Electronics 10, 2020.
Qu, Z., H. Li, Y. Wang, J. Zhang, A. Abu-Siada, and Y. Yao, "Detection of Electricity Theft Behavior Based on Improved Synthetic Minority Oversampling Technique and Random Forest Classifier", energies, 2020.
Kimura, T., "Customer churn prediction with hybrid resampling and ensembe learning", Journal of Management Information and Decision Sciences, 25, 1-23, 2022.
Otoo, G., Analysis of credit card fraud detection methods. 2021, Ashesi University College. p. 48
Batista, G.E.a.P.A., R.C. Prati, and M.C. Monard, "A study of the behavior of several methods for balancing machine learning training data", SIGKDD Explorations, 6, 20-29, 2003.
Khleel, N.a.A. and K. Nehéz, "A novel approach for software defect prediction using CNN and GRU based on SMOTE Tomek method", Journal of Intelligent Information Systems (2023) 60:673–707, 60, 673-707, 2023.
Luthra, R., G. Nath, and R. Chellani, "A review on class imbalanced correction techniques: A case of credit card default prediction on a highly imbalanced dataset", Praxis Business School, 2019.
Zhao, Z., T. Cui, S. Ding, J. Li, and A.G. Bellotti, "Resampling Techniques Study on Class Imbalance Problem in Credit Risk Prediction", Mathematics, 12, 2024.
Namvar, A., M. Siami, F. Rabhi, and M. Naderpour, "Credit risk prediction in an imbalanced social lending environment", International Journal of Computational Intelligence Systems,, 11, 925-935, 2018.
Chowdhury, S. and M.P. Schoen, "Research Paper Classification using Supervised Machine Learning Techniques", Intermountain Engineering, Technology and Computing (IETC), 2020.
Kumar, V., G.S. Lalotra, P. Sasikala, D.S. Rajput, R. Kaluri, K. Lakshmanna, M. Shorfuzzaman, A. Alsufyani, and M. Uddin, "Addressing Binary Classification over Class Imbalanced Clinical Datasets Using Computationally Intelligent Techniques", healthcare, 10, 2022.
Juma, S.W. (2021) Robust statistical learning for optimal classification of imbalanced data. 1-92.
Rubaidi, Z.S., B.B. Ammar, and M.B. Aouicha, "Fraud Detection Using Large-scale Imbalance Dataset", International Journal on Artificial Intelligence Tools, 31, 2-23, 2022.
Milli, M.E.F., I.D. Kocakoc, and S. Aras, "Investigating the Effect of Class Balancing Methods on the Performance of Machine Learning Techniques: Credit Risk Application", izmir journal of management, 5, 55-69, 2024.
Ali, M.R., Prediction Accuracy of Financial Data-Applying Several Resampling Techniques, in Computer Science 2020, North Dakota State University of Agriculture and Applied Science Fargo, North Dakota. p. 36
Pamuk, M., R.O. Grendel, and M. Schumann, "Towards ML-based Platforms in Finance Industry – An ML Approach to Gener o Generate Corporate Bankruptcy Probabilities based on Annual Financial Statements ", AIS Electronic Library (AISeL), 8, 2021.
DOI: https://doi.org/10.47738/jads.v6i3.713
Refbacks
- There are currently no refbacks.

Journal of Applied Data Sciences
| ISSN | : | 2723-6471 (Online) |
| Publisher | : | Bright Publisher |
| Website | : | http://bright-journal.org/JADS |
| : | taqwa@amikompurwokerto.ac.id (principal contact) | |
| support@bright-journal.org (technical issues) |
This work is licensed under a Creative Commons Attribution-ShareAlike 4.0




.png)