Applied Density-Based Clustering Techniques for Classifying High-Risk Customers: A Case Study of Commercial Banks in Vietnam

Nguyen Minh Nhat

Abstract


Understanding and effectively engaging with customers is paramount in today's rapidly evolving business landscape. With rapid technological advances, banks have unprecedented opportunities to improve their approach to customer segmentation. This change is driven by integrating resource planning systems and digital tools, enabling a more comprehensive and data-driven understanding of customer behavior. Therefore, the study aims to evaluate the performance of various density-based clustering algorithms in classifying customers at risk of default. The algorithms analyzed include K-Means, DBSCAN, HDBSCAN, and Birch, each offering unique strengths in handling diverse data structures. Using a dataset of 77,272 customers from Vietnamese commercial banks spanning 2010 to 2022, the study rigorously assesses these models based on seven critical metrics: Davies-Bouldin Index, Silhouette Score, Adjusted Rand Index, Homogeneity, Completeness, V-Measure, and Accuracy. The results indicate that density-based methods, particularly DBSCAN and HDBSCAN, excel in identifying high-risk clusters despite challenges in cluster separation and alignment with accurate data distributions. Birch demonstrates superior cluster separation and compactness but requires further refinement for optimal accuracy. The findings underscore the potential of integrating clustering methods into credit risk management frameworks, enhancing financial institutions' predictive accuracy and operational efficiency. This research contributes to the ongoing discourse on practical credit risk assessment tools, providing valuable insights for practitioners in the banking sector. Finally, once segments are identified, banks can tailor marketing messages, product offerings, and customer experiences to better suit each group. This can lead to reduced risk, improved customer satisfaction, higher conversion rates, and ultimately increased revenue and customer segmentation in the context of technology trends is becoming an indispensable part of modern business strategy

Keywords


Credit risk management, clustering algorithms, customer risk assessment, irregular clusters, density-based clustering

Full Text:

PDF

References


M. Ahmed, R. Seraj, and S. M. S. Islam, "The k-means algorithm: A comprehensive survey and performance evaluation", Electronics, vol. 9, no. 8, pp. 1-12, 2020. https://doi.org/10.3390/electronics9081295.

M. Ala'raj and M. F. Abbod, "A new hybrid ensemble credit scoring model based on classifiers consensus system approach", Expert Systems with Applications, vol. 64, no. 1, pp. 36-55, 2016. https://doi.org/10.1016/j.eswa.2016.07.017.

M. Ala'raj and M. F. Abbod, "Classifiers consensus system approach for credit scoring", Knowledge-Based Systems, vol. 104, no. 7, pp. 89-105, 2016. https://doi.org/10.1016/j.knosys.2016.04.013.

E. Amigó, J. Gonzalo, J. Artiles, and F. Verdejo, "A comparison of extrinsic clustering evaluation metrics based on formal constraints", Information Retrieval, vol. 12, no. 7, pp. 461-486, 2009. https://doi.org/10.1007/s10791-008-9066-8.

J. N. Crook, D. B. Edelman, and L. C. Thomas, "Recent developments in consumer credit risk assessment", European Journal of Operational Research, vol. 183, no. 3, pp. 1447-1465, 2007. https://doi.org/10.1016/j.ejor.2006.09.100.

M. Halkidi, Y. Batistakis, and M. Vazirgiannis, "On clustering validation techniques", Journal of Intelligent Information Systems, vol. 17, no. 1, pp. 107-145, 2001. https://doi.org/10.1023/A:1012801612483.

T. Hosaka, "Bankruptcy prediction using imaged financial ratios and convolutional neural networks", Expert Systems with Applications, vol. 117, no. 3, pp. 287-299, 2019. https://doi.org/10.1016/j.eswa.2018.09.039.

G. Kou, Y. Peng, and G. Wang, "Evaluation of clustering algorithms for financial risk analysis using MCDM methods", Information Sciences, vol. 275, no. 8, pp. 1-12, 2014. https://doi.org/10.1016/j.ins.2014.02.137.

P. R. Kumar and V. Ravi, "Bankruptcy prediction in banks and firms via statistical and intelligent techniques - A review", European Journal of Operational Research, vol. 180, no. 1, pp. 1-28, 2007. https://doi.org/10.1016/j.ejor.2006.08.043.

X. Li, P. Zhang, and G. Zhu, "DBSCAN clustering algorithms for non-uniform density data and its application in urban rail passenger aggregation distribution", Energies, vol. 12, no. 19, pp. 1-22, 2019. https://doi.org/10.3390/en12193722.

M. Moscatelli, F. Parlapiano, S. Narizzano, and G. Viggiano, "Corporate default forecasting with machine learning", Expert Systems with Applications, vol. 161, no. 1, pp. 1-17, 2020. https://doi.org/10.1016/j.eswa.2020.113567.

I. K. Nti, A. F. Adekoya, and B. A. Weyori, "A systematic review of fundamental and technical analysis of stock market predictions", Artificial Intelligence Review, vol. 53, no. 4, pp. 3007-3057, 2020. https://doi.org/10.1007/s10462-019-09754-z.

D. L. Olson, D. Delen, and Y. Meng, "Comparative analysis of data mining methods for bankruptcy prediction", Decision Support Systems, vol. 52, no. 2, pp. 464-473, 2012. https://doi.org/10.1016/j.dss.2011.10.007.

Y. Qu, P. Quan, M. Lei, and Y. Shi, "Review of bankruptcy prediction using machine learning and deep learning techniques", Procedia Computer Science, vol. 162, pp. 895-899, 2019. https://doi.org/10.1016/j.procs.2019.12.065.

P. J. Rousseeuw, "Silhouettes: A graphical aid to the interpretation and validation of cluster analysis", Journal of Computational and Applied Mathematics, vol. 20, pp. 53-65, 1987. https://doi.org/10.1016/0377-0427(87)90125-7.

S. D. Vrontos, J. Galakis, and I. D. Vrontos, "Modeling and predicting US recessions using machine learning techniques", International Journal of Forecasting, vol. 37, no. 2, pp. 647-671, 2021. https://doi.org/10.1016/j.ijforecast.2020.08.005.

L. Wang, P. Chen, L. Chen, and J. Mou, "Ship AIS trajectory clustering: An HDBSCAN-based approach", Journal of Marine Science and Engineering, vol. 9, no. 6, pp. 1-20, 2021. https://doi.org/10.3390/jmse9060566.

R. Xu and D. Wunsch, "Survey of clustering algorithms", IEEE Transactions on Neural Networks, vol. 16, no. 3, pp. 645-678, 2005. https://doi.org/10.1109/TNN.2005.845141.

B. W. Yap, S. H. Ong, and N. H. M. Husain, "Using data mining to improve assessment of credit worthiness via credit scoring models", Expert Systems with Applications, vol. 38, no. 10, pp. 13274-13283, 2011. https://doi.org/10.1016/j.eswa.2011.04.147.

T. Zhang, R. Ramakrishnan, and M. Livny, "BIRCH: An efficient data clustering method for very large databases", ACM Sigmod Record, vol. 25, no. 2, pp. 103-114, 1996. https://doi.org/10.1145/235968.233324.

F. Yang and S. Gu, "Industry 4.0, a revolution that requires technology and national strategies", Complex & Intelligent Systems, vol. 7, no. 3, pp. 1311-1325, 2021. https://doi.org/10.1007/s40747-020-00267-9.

M. Westerdijk, J. Zuurbier, M. Ludwig, and S. Prins, "Defining care products to finance health care in the Netherlands", The European Journal of Health Economics, vol. 13, no. 2, pp. 203-221, 2012. https://doi.org/10.1007/s10198-011-0302-6.

D. Tingley, J. Ásmundsson, E. Borodzicz, A. Conides, B. Drakeford, I. R. Eðvarðsson, D. Holm, K. Kapiris, S. Kuikka, and B. Mortensen, "Risk identification and perception in the fisheries sector: Comparisons between the Faroes, Greece, Iceland and UK", Marine Policy, vol. 34, no. 6, pp. 1249-1260, 2010. https://doi.org/10.1016/j.marpol.2010.05.002.

Y. Thakare and S. Bagal, "Performance evaluation of K-means clustering algorithm with various distance metrics", International Journal of Computer Applications, vol. 110, no. 11, pp. 12-16, 2015. https://doi.org/10.5120/19360-0929.

V. Singhal, A. B. Singh, V. Ahuja, and R. Gera, "Consumer segmentation in the fashion industry using social media: An empirical analysis", Journal of Information and Organizational Sciences, vol. 47, no. 2, pp. 399-419, 2023. https://doi.org/10.31341/jios.47.2.9.

A. Sheikh, T. Ghanbarpour, and D. Gholamiangonabadi, "A preliminary study of fintech industry: A two-stage clustering analysis for customer segmentation in the B2B setting", Journal of Business-to-Business Marketing, vol. 26, no. 2, pp. 197-207, 2019. https://doi.org/10.1080/1051712X.2019.1603420.

L. Rivera, D. Gligor, and Y. Sheffi, "The benefits of logistics clustering", International Journal of Physical Distribution & Logistics Management, vol. 46, no. 3, pp. 242-268, 2016. https://doi.org/10.1108/IJPDLM-10-2014-0243.

M. Argüelles, C. Benavides, and I. Fernández, "A new approach to the identification of regional clusters: Hierarchical clustering on principal components", Applied Economics, vol. 46, no. 21, pp. 1-19, 2014. https://doi.org/10.1080/00036846.2014.904491.

M. Bagherzadeh, M. Ghaderi, and A. S. Fernandez, "Coopetition for innovation – The more, the better? An empirical study based on preference disaggregation analysis", European Journal of Operational Research, vol. 297, no. 2, pp. 695-708, 2022. https://doi.org/10.1016/j.ejor.2021.06.010.

G. G. Dureti, M. P. Tabe-Ojong, and E. Owusu-Sekyere, "The new normal? Cluster farming and smallholder commercialization in Ethiopia", Agricultural Economics, vol. 54, no. 6, pp. 900-920, 2023. https://doi.org/10.1111/agec.12790.

W. Hadhri, R. Arvanitis, and H. M'Henni, "Determinants of innovation activities in small and open economies: The Lebanese business sector", Journal of Innovation Economics & Management, vol. 21, no. 3, pp. 77-107, 2016. https://doi.org/10.3917/jie.021.0077.

W. S. Hwang and H. S. Kim, "Does the adoption of emerging technologies improve technical efficiency? Evidence from Korean manufacturing SMEs", Small Business Economics, vol. 59, no. 2, pp. 627-643, 2022. https://doi.org/10.1007/s11187-021-00554-w.

G. J. Oyewole and G. A. Thopil, "Data clustering: Application and trends", Artificial Intelligence Review, vol. 56, no. 7, pp. 6439-6475, 2023. https://doi.org/10.1007/s10462-022-10325-y




DOI: https://doi.org/10.47738/jads.v5i4.344

Refbacks

  • There are currently no refbacks.



Barcode

Journal of Applied Data Sciences

ISSN:2723-6471 (Online)
Publisher:Bright Publisher
Website:http://bright-journal.org/JADS
Email:taqwa@amikompurwokerto.ac.id (principal contact)
  support@bright-journal.org (technical issues)

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0