Optimizing Stunting Detection through SMOTE and Machine Learning: a Comparative Study of XGBoost, Random Forest, SVM, and k-NN
Abstract
Stunting is a vital public health priority that affects millions of children from all over the world, especially in developing countries, where chronic malnutrition impairs their physical growth and cognitive development. Early detection of stunting is necessary for its timely intervention to reduce long-lasting effects. The following study deals with the application of higher-end machine learning techniques in order to detect stunting with more accuracy, using XGBoost, Random Forest, SVM, and k-NN algorithms. Using a dataset sourced from Kaggle, containing 10,000 samples of anthropometric and demographic features, we addressed the significant class imbalance of the data; the number of samples representing stunted children was only 15% of the total. We surmounted this limitation using SMOTE to generate synthetic data in order to balance the representation for this minority class. Further feature selection to improve the performance and interpretability of the model was done using backward elimination, where less impactful features like "Body Length" and "Breastfeeding" were systematically excluded, while putting more emphasis on more predictive variables such as weight, age, and socio-economic indicators. The evaluation of machine learning models showed significant improvements in performance with the integration of SMOTE and optimized feature selection, especially regarding recall and ROC-AUC metrics, which are critical in healthcare settings where the minimization of false negatives is of high importance. XGBoost was the best-performing model among those evaluated, yielding an accuracy of 0.8574, a recall of 0.8914, and an ROC-AUC of 0.9311, hence balancing precision and sensitivity more appropriately than other models. These results emphasize the efficiency of XGBoost in stunting detection while overcoming challenges arising from imbalanced datasets. It then illustrates the potential of merging machine learning techniques with synthetic data augmentation methodologies for the optimization of outcomes related to population health, and forms a basis for healthcare practitioners and policymakers by locating the at-risk children on time. The findings not only point to the importance of advanced data-driven approaches in stunting detection but also lay the ground for future research on machine learning applications in the fight against other malnutrition-related public health challenges, which could be crucial for improving child health and well-being across the world.
Keywords
Full Text:
PDFReferences
E. H. Nugrahani, S. Nurdiati, F. Bukhari, M. K. Najib, D. M. Sebastian, and P. A. N. Fallahi, “Sensitivity and feature importance of climate factors for predicting fire hotspots using machine learning methods,” IAES International Journal of Artificial Intelligence, vol. 13, no. 2, pp. 2210–2223, Jun. 2024, doi: 10.11591/ijai.v13.i2.pp2212-2225.
F. Hilali Moh’d, K. Anwar Notodiputro, and Y. Angraini, “Enhancing interpretability in random forest: Leveraging inTrees for association rule extraction insights,” IAES International Journal of Artificial Intelligence (IJ-AI), vol. 13, no. 4, p. 4054, Dec. 2024, doi: 10.11591/ijai.v13.i4.pp4054-4061.
C. Angelica, C. Charleen, and A. Wibowo, “Elevating fraud detection: machine learning models with computational intelligence optimization,” IAES International Journal of Artificial Intelligence (IJ-AI), vol. 13, no. 4, p. 4273, Dec. 2024, doi: 10.11591/ijai.v13.i4.pp4273-4280.
D. A. Kristiyanti, S. A. Sanjaya, V. C. Tjokro, and J. Suhali, “Dealing imbalance dataset problem in sentiment analysis of recession in Indonesia,” IAES International Journal of Artificial Intelligence, vol. 13, no. 2, pp. 2058–2070, Jun. 2024, doi: 10.11591/ijai.v13.i2.pp2060-2072.
S. Makubhai, G. R. Pathak, and P. R. Chandre, “Comparative analysis of explainable artificial intelligence models for predicting lung cancer using diverse datasets,” IAES International Journal of Artificial Intelligence, vol. 13, no. 2, pp. 1978–1989, Jun. 2024, doi: 10.11591/ijai.v13.i2.pp1980-1991.
S. H, “Multi-Label Feature Aware XGBoost Model For Student Performance Assessment Using Behavior Data in Online Learning Environment,” IAES International Journal of Artificial Intelligence (IJ-AI), vol. 13, no. 4, p. 4537, Dec. 2024, doi: 10.11591/ijai.v13.i4.pp4537-4543.
W. Chimphlee and S. Chimphlee, “Hyperparameters optimization XGBoost for network intrusion detection using CSE-CIC-IDS 2018 dataset,” IAES International Journal of Artificial Intelligence, vol. 13, no. 1, pp. 817–826, Mar. 2024, doi: 10.11591/ijai.v13.i1.pp817-826.
N. Effendy, M. Z. A. Fadhilah, D. W. Kraton, and H. A. Abrar, “The prediction of thermal sensation in building using support vector machine and extreme gradient boosting,” IAES International Journal of Artificial Intelligence, vol. 13, no. 3, pp. 2963–2970, Sep. 2024, doi: 10.11591/ijai.v13.i3.pp2963-2970.
A. Febriani, “Improved Hybrid Machine and Deep Learning Model for Optimization of Smart Egg Incubator,” Journal of Applied Data Sciences, vol. 5, no. 3, pp. 1052–1068, Sep. 2024, doi: 10.47738/jads.v5i3.304.
R. Govindarajan, V. Balaji, J. Arumugam, T. A. Assegie, and R. Mothukuri, “Evaluation of sequential feature selection in improving the K-nearest neighbor classifier for diabetes prediction,” IAES International Journal of Artificial Intelligence, vol. 13, no. 2, pp. 1567–1573, Jun. 2024, doi: 10.11591/ijai.v13.i2.pp1567-1573.
I. Slamet, “Retinopathy Classification using Convolutional Neural Network Method with Adaptive Momentum Optimization and Applied Batch Normalization,” Journal of Applied Data Sciences, vol. 5, no. 3, pp. 1123–1133, Sep. 2024, doi: 10.47738/jads.v5i3.309.
M. Ipa, Y. Yuliasih, E. P. Astuti, A. D. Laksono, and W. Ridwan, “STAKEHOLDERS’ ROLE IN THE IMPLEMENTATION OF STUNTING MANAGEMENT POLICIES IN GARUT REGENCY,” Jurnal Administrasi Kesehatan Indonesia, vol. 11, no. 1, pp. 26–35, Jun. 2023, doi: 10.20473/jaki.v11i1.2023.26-35.
S. Armoogum, “Breast Cancer Prediction Using Metrics-Based Classification,” Journal of Applied Data Sciences, vol. 5, no. 3, pp. 1508–1519, Sep. 2024, doi: 10.47738/jads.v5i3.351.
S. Munawaroh, M. N. Fajri, and S. R. Ajija, “THE EFFECTS OF SOCIAL ASSISTANCE PROGRAMS ON STUNTING PREVALENCE RATES IN INDONESIA,” Indonesian Journal of Health Administration, vol. 12, no. 1, pp. 74–85, Jun. 2024, doi: 10.20473/jaki.v12i1.2024.74-85.
I. E. Purba, Y. G. Tarigan, A. Zendrato, A. Purba, and T. Sinaga, “Maternal factors associated with stunting among children under two years in South Nias, Indonesia: a cross-sectional study,” International Journal of Public Health Science (IJPHS), vol. 13, no. 3, p. 1349, Sep. 2024, doi: 10.11591/ijphs.v13i3.24316.
R. Tanjung, D. Lestrina, and J. Sinaga, “Spatial analysis of environmental sanitation and stunting incidents,” International Journal of Public Health Science (IJPHS), vol. 13, no. 4, p. 1968, Dec. 2024, doi: 10.11591/ijphs.v13i4.23442.
M. K. Anam, S. Defit, Haviluddin, L. Efrizoni, and M. B. Firdaus, “Early Stopping on CNN-LSTM Development to Improve Classification Performance,” Journal of Applied Data Sciences, vol. 5, no. 3, pp. 1175–1188, Sep. 2024, doi: 10.47738/jads.v5i3.312.
R. K. Sari, C. P. Mayangsari, I. D. Mashoedi, Y. S. N. Intan, S. Trisnadi, and D. F. Aprilyanti, “Strengthening emotional intelligence intervention on behavior changes of mothers in stunting prevention,” International Journal of Public Health Science (IJPHS), vol. 13, no. 2, p. 536, Jun. 2024, doi: 10.11591/ijphs.v13i2.23652.
H. A. Abdelhafez and A. A. Amer, “Machine Learning Techniques for Diabetes Prediction: A Comparative Analysis,” Journal of Applied Data Sciences, vol. 5, no. 2, pp. 792–807, May 2024, doi: 10.47738/jads.v5i2.219.
A. R. Hananto, “Identifying Student Learning Styles Using Support Vector Machine in Felder-Silverman Model,” Journal of Applied Data Sciences, vol. 5, no. 3, pp. 1495–1507, Sep. 2024, doi: 10.47738/jads.v5i3.337.
A. Suryaputra Paramita, I. Maryati, and L. M. Tjahjono, “Implementation of the K-Nearest Neighbor Algorithm for the Classification of Student Thesis Subjects,” Journal of Applied Data Sciences, vol. 3, no. 3, pp. 128–136, 2022.
N. Trianasari and T. A. Permadi, “Analysis of Product Recommendation Models at Each Fixed Broadband Sales Location Using K-Means, DBSCAN, Hierarchical Clustering, SVM, RF, and ANN,” Journal of Applied Data Sciences, vol. 5, no. 2, pp. 636–652, May 2024, doi: 10.47738/jads.v5i2.210.
A. Azis, “Application of the Vector Machine Support Method in Twitter Social Media Sentiment Analysis Regarding the Covid-19 Vaccine Issue in Indonesiaid 2 * corresponding author,” Journal of Applied Data Sciences, vol. 2, no. 3, pp. 102–108, 2021.
P. S. Siregar, R. G. Hatika, and B. H. Hayadi, “Multiple Choice Question Difficulty Level Classification with Multi Class Confusion Matrix in the Online Question Bank of Education Gallery,” Journal of Applied Data Sciences, vol. 4, no. 4, pp. 392–406, Dec. 2023, doi: 10.47738/jads.v4i4.132.
L. Jen and Y.-H. Lin, “A Brief Overview of the Accuracy of Classification Algorithms for Data Prediction in Machine Learning Applicationstw 2 * corresponding author,” Journal of Applied Data Sciences, vol. 2, no. 3, pp. 84–92, 2021.
DOI: https://doi.org/10.47738/jads.v6i1.494
Refbacks
- There are currently no refbacks.

Journal of Applied Data Sciences
| ISSN | : | 2723-6471 (Online) |
| Publisher | : | Bright Publisher |
| Website | : | http://bright-journal.org/JADS |
| : | taqwa@amikompurwokerto.ac.id (principal contact) | |
| support@bright-journal.org (technical issues) |
This work is licensed under a Creative Commons Attribution-ShareAlike 4.0




.png)