Hybrid Ensemble Learning with SMOTEENN and Soft Voting for Stunting Risk Prediction: A SHAP-Based Explainable Approach

Nuwairy El Furqany, Muhammad Subianto, Asep Rusyana

Abstract


Stunting remains a critical public health concern in Indonesia, with long-term consequences for physical growth, cognitive development, and human capital. This study introduces a hybrid machine learning framework to predict household-level stunting risk by integrating Synthetic Minority Over-sampling Technique with Edited Nearest Neighbors (SMOTEENN), soft voting ensemble, and SHapley Additive exPlanations (SHAP). The objective is to enhance both predictive accuracy and interpretability in identifying high-risk households. A dataset of 115,579 household records from West Sumatra, comprising 20 demographic, socioeconomic, health, and housing predictors, was utilized. Preprocessing steps included handling missing values, categorical encoding, and applying SMOTEENN exclusively on the training set to mitigate class imbalance. The baseline models demonstrated limited sensitivity, with XGBoost performing best at 74.56% accuracy and 71.08% F1-score on imbalanced data. After applying SMOTEENN, performance improved substantially, with XGBoost achieving 91.82% accuracy and 91.74% F1-score. Further improvements were obtained through hybridization, where the Random Forest and XGBoost soft voting ensemble reached 91.95% accuracy and 92.46% F1-score, representing a notable gain over individual classifiers. SHAP analysis added interpretability by identifying family members, education level, diverse food consumption, occupation, and drinking water source as dominant predictors of stunting risk. The novelty of this study lies in the integration of SMOTEENN with ensemble learning and SHAP, providing not only robust performance but also transparency in feature contributions. The findings demonstrate that the proposed framework improves sensitivity to minority classes, delivers superior predictive accuracy compared to baseline models, and offers interpretable insights to guide targeted interventions. By combining methodological rigor with explainability, this research contributes a practical decision-support tool for policymakers, supporting early detection of at-risk households and accelerating stunting reduction efforts in Indonesia.


Keywords


Stunting Prediction; Hybrid Machine Learning; SMOTEENN; Soft Voting Ensemble; SHAP Interpretability; Class Imbalance; Public Health; Stunting

Full Text:

PDF

References


Kementerian Kesehatan Republik Indonesia (Kemenkes RI), Survei Status Gizi Indonesia (SSGI) Tahun 2024 [Indonesia Nutritional Status Survey 2024], Jakarta, Indonesia: Badan Kebijakan Pembangunan Kesehatan, 2024. [Online]. Available: https://www.badankebijakan.kemkes.go.id/survei-status-gizi-indonesia-ssgi-2024. [Accessed: Aug. 1, 2025].

A. T. Mulyani, M. A. Khairinisa, A. Khatib, and A. Y. Chaerunisaa, "Understanding Stunting: Impact, Causes, and Strategy to Accelerate Stunting Reduction—A Narrative Review," Nutrients, vol. 17, no. 9, p. 1493, 2025. doi: 10.3390/nu17091493.

G. Nduwayezu, P. Zhao, P. Pilesjö, J. P. Bizimana, and A. Mansourian, "Multilevel small-area childhood stunting risk estimation: Insights from spatial ensemble learning, agro-ecological and environmentally remotely sensed indicators," Environmental and Sustainability Indicators, vol. 27, p. 100822, 2025. doi: 10.1016/j.indic.2025.100822.

M. Ayele, G. A. Baye, S. H. Yesuf, et al., "Predicting stunting status among under-five children in Ethiopia using ensemble machine learning algorithms," Scientific Reports, vol. 15, p. 27907, 2025. doi: 10.1038/s41598-025-03206-1.

A. Fernández, S. García, M. Galar, R. C. Prati, B. Krawczyk, and F. Herrera, Learning from Imbalanced Data Sets, Cham, Switzerland: Springer International Publishing, 2018. doi: 10.1007/978-3-319-98074-4.

G. E. A. P. A. Batista, R. C. Prati, and M. C. Monard, "A study of the behavior of several methods for balancing machine learning training data," SIGKDD Explorations, vol. 6, no. 1, pp. 20–29, Jun. 2004. doi: 10.1145/1007730.1007735.

L. Siena, T. H. Saragih, R. A. Nugroho, D. Kartini, Muliadi, and W. Caesarendra, "Evaluation of the Impact of SMOTEENN on Monkeypox Case Classification Performance Using Boosting Algorithms," Indonesian Journal of Electronics, Electromedical Engineering, and Medical Informatics, vol. 7, no. 2, pp. 203–220, 2025. doi: 10.35882/ijeeemi.v7i2.77.

T. G. Dietterich, "Ensemble methods in machine learning," in Multiple Classifier Systems, G. Goos, J. Hartmanis, and J. Van Leeuwen, Eds., Berlin, Heidelberg: Springer, 2000, vol. 1857, pp. 1–15. doi: 10.1007/3-540-45014-9_1.

M. Nguyen, K. Cao-Van, L. G. Minh, T. X. Bui, and S. Hong, "Hybrid Machine Learning Models Using Soft Voting Classifier for Financial Distress Prediction," SSRN Electronic Journal, Aug. 2024. doi: 10.2139/ssrn.4941751.

N. Gadde, A. Mohapatra, D. Tallapragada, K. Mody, N. Vijay, and A. Gottumukhala, "Explainable AI for dynamic ensemble models in high-stakes decision-making," International Journal of Science and Research Archive, vol. 13, no. 2, pp. 1170–1176, 2024. doi: 10.30574/ijsra.2024.13.2.2091.

S. Ndagijimana, I. H. Kabano, E. Masabo, and J. M. Ntaganda, "Prediction of stunting among under-5 children in Rwanda using machine learning techniques," Journal of Preventive Medicine & Public Health, vol. 56, no. 1, pp. 41–49, Jan. 2023. doi: 10.3961/jpmph.22.388.

H. Shen, H. Zhao, and Y. Jiang, "Machine learning algorithms for predicting stunting among under-five children in Papua New Guinea," Children, vol. 10, no. 10, p. 1638, 2023. doi: 10.3390/children10101638.

T. Sugihartono, B. Wijaya, Marini, A. F. Alkayes, and H. A. Anugrah, "Optimizing stunting detection through SMOTE and machine learning: A comparative study of XGBoost, Random Forest, SVM, and k-NN," Journal of Applied Data Sciences, vol. 6, no. 1, pp. 667–682, 2024. doi: 10.47738/jads.v6i1.494.

A. B. Zemariam, B. B. Abate, A. W. Alamaw, E. S. Lake, G. Yilak, M. Ayele, B. D. Tilahun, and H. S. Ngusie, "Prediction of stunting and its socioeconomic determinants among adolescent girls in Ethiopia using machine learning algorithms," PLoS One, vol. 20, no. 1, p. e0316452, Jan. 2025. doi: 10.1371/journal.pone.0316452.

Badan Kependudukan dan Keluarga Berencana Nasional Republik Indonesia (BKKBN RI), Publikasi Data Keluarga Berisiko Stunting Tahun 2024 [Publication of Families at Risk of Stunting Data 2024], Jakarta, Indonesia: Direktorat Pelaporan dan Statistik, 2024.

A. Tawakuli and T. Engel, "Make your data fair: A survey of data preprocessing techniques that address biases in data towards fair AI," Journal of Engineering Research, 2024. doi: 10.1016/j.jer.2024.06.016.

F. Bolikulov, R. Nasimov, A. Rashidov, F. Akhmedov, and Y.-I. Cho, "Effective methods of categorical data encoding for artificial intelligence algorithms," Mathematics, vol. 12, no. 16, p. 2553, 2024. doi: 10.3390/math12162553.

N. V. Chawla, K. W. Bowyer, L. O. Hall, and W. P. Kegelmeyer, "SMOTE: Synthetic minority over-sampling technique," Journal of Artificial Intelligence Research, vol. 16, pp. 321–357, 2002. doi: 10.1613/jair.953.

G. Husain, D. Nasef, R. Jose, J. Mayer, M. Bekbolatova, T. Devine, and M. Toma, "SMOTE vs. SMOTEENN: A study on the performance of resampling algorithms for addressing class imbalance in regression models," Algorithms, vol. 18, no. 1, p. 37, 2025. doi: 10.3390/a18010037.

D. W. Hosmer and S. Lemeshow, Applied Logistic Regression, 2nd ed., New York, NY, USA: John Wiley & Sons, 2000. doi: 10.1002/0471722146.

E. Métais, F. Meziane, M. Saraee, V. Sugumaran, and S. Vadera, Eds., Natural Language Processing and Information Systems: 21st International Conference on Applications of Natural Language to Information Systems (NLDB 2016), Salford, UK, June 22–24, 2016, Proceedings, vol. 9612, Cham, Switzerland: Springer International Publishing, 2016. doi: 10.1007/978-3-319-41754-7.

R. G. Brereton and G. R. Lloyd, "Support vector machines for classification and regression," The Analyst, vol. 135, no. 2, pp. 230–267, 2010. doi: 10.1039/B918972F.

T. Chen and C. Guestrin, "XGBoost: A scalable tree boosting system," in Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD ’16), San Francisco, CA, USA, Aug. 2016, pp. 785–794. doi: 10.1145/2939672.2939785.

P. Chithuloori and J. M. Kim, "Soft voting ensemble classifier for liquefaction prediction based on SPT data," Artificial Intelligence Review, vol. 58, p. 228, 2025. doi: 10.1007/s10462-025-11230-w.

P. Mahajan, S. Uddin, F. Hajati, and M. A. Moni, "Ensemble learning for disease prediction: A review," Healthcare, vol. 11, no. 12, p. 1808, 2023. doi: 10.3390/healthcare11121808.

K. Cao-Van, T. Cao Minh, L. Gia Minh, T. B. Q. Thi, and H. Minh Tan, "Soft-voting ensemble model: An efficient learning approach for predictive prostate cancer risk," Vietnam Journal of Computer Science, vol. 11, no. 4, 2024. doi: 10.1142/S2196888824500155.

J. Han and M. Kamber, Data Mining: Concepts and Techniques, 3rd ed., Burlington, MA, USA: Morgan Kaufmann Publishers, 2012. doi: 10.1016/C2009-0-61819-5.

S. M. Lundberg and S.-I. Lee, "A unified approach to interpreting model predictions," in Advances in Neural Information Processing Systems (NeurIPS 30), Red Hook, NY, USA: Curran Associates, Inc., 2017, pp. 4765–477.

C. Molnar, Interpretable Machine Learning: A Guide for Making Black Box Models Explainable, 2nd ed., 2020.

A. S. Antonini, J. Tanzola, L. Asiain, G. R. Ferracutti, S. M. Castro, E. A. Bjerg, and M. L. Ganuza, "Machine Learning model interpretability using SHAP values: Application to Igneous Rock Classification task," Applied Computing and Geosciences, vol. 23, p. 100178, 2024. doi: 10.1016/j.acags.2024.100178.




DOI: https://doi.org/10.47738/jads.v6i4.829

Refbacks

  • There are currently no refbacks.



Barcode

Journal of Applied Data Sciences

ISSN:2723-6471 (Online)
Publisher:Bright Publisher
Website:http://bright-journal.org/JADS
Email:taqwa@amikompurwokerto.ac.id (principal contact)
  support@bright-journal.org (technical issues)

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0