Beyond Accuracy: A Data-Validity and Target-Reconstruction Audit of Machine Learning for Stunting Classification in 40,071 Indonesian Toddlers
DOI:
https://doi.org/10.56705/qyaph181Abstract
Introduction: Machine-learning models for stunting classification often report very high accuracy, but this performance may partly result from reconstructing anthropometric rules used to define the target rather than predicting future stunting risk. This study evaluated the validity of machine-learning-based stunting classification using 40,071 toddler records from Jeneponto Regency, Indonesia. Method: Raw spreadsheet integrity, annual sampling structure, anthropometric variable representation, target consistency, duplicate profiles, and feature provenance were audited before modeling. The 2024 cohort was used for primary analysis because the 2021–2023 datasets contained only stunted cases. Random Forest models were evaluated using five feature configurations under stratified and duplicate-profile-aware grouped cross-validation, supported by cross-algorithm robustness analysis. Results: The categorical stunting label showed 99.9937% agreement with the WHO HAZ < −2 criterion. Under grouped validation, Age, Gender, and Height achieved a ROC-AUC of 0.9886, which decreased to 0.7756 after Height was removed. Adding Weight-for-Height Z-score increased ROC-AUC to 0.9970, indicating indirect reintroduction of height information. Furthermore, 83.56% of classification errors occurred within ±0.20 SD of the WHO HAZ threshold, with 30.74-fold higher odds of misclassification near this boundary. Conclusion: High machine-learning performance in this dataset primarily reflects contemporaneous anthropometric target reconstruction. Stunting studies should therefore distinguish automated classification from prospective prediction and explicitly evaluate data provenance, feature construction, and validation design.
References
[1] World Health Organization, United Nations Children's Fund, and World Bank Group, Levels and Trends in Child Malnutrition: UNICEF/WHO/World Bank Group Joint Child Malnutrition Estimates—Key Findings of the 2025 Edition. Geneva, Switzerland: World Health Organization, 2025.
[2] World Health Organization, WHO Child Growth Standards: Length/Height-for-Age, Weight-for-Age, Weight-for-Length, Weight-for-Height and Body Mass Index-for-Age: Methods and Development. Geneva, Switzerland: WHO, 2006.
[3] E. Lestari, A. Siregar, A. K. Hidayat, and A. A. Yusuf, “Stunting and its association with education and cognitive outcomes in adulthood: A longitudinal study in Indonesia,” PLOS ONE, vol. 19, no. 5, Art. no. e0295380, 2024, doi: https://doi.org/10.1371/journal.pone.0295380.
[4] M. A. Alam et al., “Impact of early-onset persistent stunting on cognitive development at 5 years of age: Results from a multi-country cohort study,” PLOS ONE, vol. 15, no. 1, Art. no. e0227839, 2020, doi: https://doi.org/10.1371/journal.pone.0227839.
[5] Badan Kebijakan Pembangunan Kesehatan, Survei Status Gizi Indonesia (SSGI) 2024 Dalam Angka. Jakarta, Indonesia: Kementerian Kesehatan Republik Indonesia, 2025.
[6] H. Anastasia et al., “Determinants of stunting in children under five years old in South Sulawesi and West Sulawesi Province: 2013 and 2018 Indonesian Basic Health Survey,” PLOS ONE, vol. 18, no. 5, Art. no. e0281962, 2023, doi: https://doi.org/10.1371/journal.pone.0281962.
[7] D. Azriani, D. Agustian, Y. Zuhairini, I. N. Yulita, and M. Dhamayanti, “Prediction models for stunting at 2-years-old from Indonesian newborn population,” BMC Pediatrics, vol. 25, Art. no. 718, 2025, doi: https://doi.org/10.1186/s12887-025-06096-4.
[8] A. H. RS, P. Purnamawati, H. Jaya, A. Arfandi, I. Suhardi, S. Widodo, and S. Lonang, “Dataset Stunting and Nutritional Status of Toddler from Jeneponto Regency, South Sulawesi, Indonesia,” Mendeley Data, Version 4, 2026, doi: https://doi.org/10.17632/wzwpc9j5bx.4.
[9] A. Wicaksono, D. Prasetyo, Y. Mar’atullatifah, D. U. Iswavigra, H. Mahmudah, and A. Hapsari, “Data Analysis and Explainable Machine Learning for Stunting Prediction,” Journal of Artificial Intelligence and Legal Technology, vol. 1, no. 1, pp. 35–44, 2025.
[10] M. Resha and A. Toding, “Predicting Independent Z-Score Stunting Through Fundamental Anthropometric Measurements Utilizing Extreme Gradient Boosting (XGBoost),” Journal Zetroem, vol. 8, no. 2, pp. 148–154, 2026, doi: https://doi.org/10.36526/ztr.v8i2.8736.
[11] H. Azis, Purnawansyah, F. Fattah, and I. P. Putri, “Performa klasifikasi K-NN dan cross-validation pada data pasien pengidap penyakit jantung,” ILKOM Jurnal Ilmiah, vol. 12, no. 2, pp. 81–86, 2020, doi: https://doi.org/10.33096/ilkom.v12i2.507.81-86.
[12] H. Azis and N. Rismayanti, “Prediksi anemia dari pixel gambar dan level hemoglobin menggunakan Random Forest Classifier,” NERO (Networking Engineering Research Operation), vol. 9, no. 1, 2024, doi: https://doi.org/10.21107/nero.v9i1.27916.
[13] F. T. Admojo and N. Rismayanti, “Estimating obesity levels using Decision Trees and K-Fold Cross-Validation: A study on eating habits and physical conditions,” Indonesian Journal of Data and Science, vol. 5, no. 1, pp. 37–44, 2024, doi: https://doi.org/10.56705/ijodas.v5i1.126.
[14] A. P. Wibowo, M. Taruk, T. E. Tarigan, and M. Habibi, “Improving mental health diagnostics through advanced algorithmic models: A case study of bipolar and depressive disorders,” Indonesian Journal of Data and Science, vol. 5, no. 1, pp. 8–14, 2024, doi: https://doi.org/10.56705/ijodas.v5i1.122.
[15] A. Halid, I. G. N. W. Arsa, R. A. Azdy, and A. A. J. Permana, “Development of a Decision Tree classifier for breast cancer diagnosis using Fine Needle Aspirate data,” Indonesian Journal of Data and Science, vol. 5, no. 3, pp. 229–236, 2024, doi: https://doi.org/10.56705/ijodas.v5i3.202.
[16] S. Kapoor and A. Narayanan, “Leakage and the reproducibility crisis in machine-learning-based science,” Patterns, vol. 4, no. 9, Art. no. 100804, 2023, doi: https://doi.org/10.1016/j.patter.2023.100804.
[17] S. E. Davis, M. E. Matheny, S. Balu, and M. P. Sendak, “A framework for understanding label leakage in machine learning for health care,” Journal of the American Medical Informatics Association, vol. 31, no. 1, pp. 274–280, 2024, doi: https://doi.org/10.1093/jamia/ocad178.
[18] D. H. Zulfikar, E. S. J. Atmadji, and B. S. W. Poetro, “Leakage-aware and explainable machine learning for healthcare claim fraud detection using imbalanced medical insurance data,” International Journal of Artificial Intelligence in Medical Issues, vol. 4, no. 1, pp. 19–34, 2026, doi: https://doi.org/10.56705/z3207345.
[19] T. A. Putra and N. Hendrastuty, “Evaluasi validitas model machine learning pada klasifikasi stunting berbasis data antropometri dan hubungan deterministik,” Building of Informatics, Technology and Science (BITS), vol. 8, no. 1, pp. 62–72, 2026, doi: https://doi.org/10.47065/bits.v8i1.9584.
[20] G. S. Collins et al., “TRIPOD+AI statement: Updated guidance for reporting clinical prediction models that use regression or machine learning methods,” BMJ, vol. 385, Art. no. e078378, 2024, doi: https://doi.org/10.1136/bmj-2023-078378.
[21] M. I. Siami, A. W. Murdiyanto, and Sumiyatun, “A comparative study of machine learning models for stress level classification using social media and lifestyle data,” International Journal of Artificial Intelligence in Medical Issues, vol. 4, no. 1, pp. 93–109, 2026, doi: https://doi.org/10.56705/r4d62a66.
[22] World Health Organization and United Nations Children's Fund, Recommendations for Data Collection, Analysis and Reporting on Anthropometric Indicators in Children Under 5 Years Old. Geneva, Switzerland: World Health Organization, 2019.
[23] L. Breiman, “Random forests,” Machine Learning, vol. 45, no. 1, pp. 5–32, 2001, doi: https://doi.org/10.1023/A:1010933404324.
[24] J. H. Friedman, “Greedy function approximation: A gradient boosting machine,” The Annals of Statistics, vol. 29, no. 5, pp. 1189–1232, 2001, doi: https://doi.org/10.1214/aos/1013203451.
[25] H. Azis, F. T. Admojo, and E. Susanti, “Analisis perbandingan performa metode klasifikasi pada dataset multiclass citra busur panah,” Techno.Com, vol. 19, no. 3, pp. 286–294, 2020, doi: https://doi.org/10.33633/tc.v19i3.3646.
[26] H. Azis and S. R. Jabir, “Chemical composition and aroma profiling: Decision Tree modeling of formalin tofu,” Journal of Embedded Systems, Security and Intelligent Systems, vol. 4, no. 2, pp. 206–211, 2023, doi: https://doi.org/10.59562/jessi.v4i2.1162.
[27] H. Azis, Y. Salim, and Y. Puspitasari, “Deep learning approaches for soft tissue classification of cephalometric images in orthodontics,” JOIV: International Journal on Informatics Visualization, vol. 10, no. 3, pp. 988–995, 2026, doi: https://doi.org/10.62527/joiv.10.3.3885.
[28] T. Saito and M. Rehmsmeier, “The precision-recall plot is more informative than the ROC plot when evaluating binary classifiers on imbalanced datasets,” PLOS ONE, vol. 10, no. 3, Art. no. e0118432, 2015, doi: https://doi.org/10.1371/journal.pone.0118432.
[29] F. Pedregosa et al., “Scikit-learn: Machine learning in Python,” Journal of Machine Learning Research, vol. 12, no. 85, pp. 2825–2830, 2011.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Nusantara Journal of Multidisciplinary Informatics
Creative Commons Attribution-NonCommercial 4.0 International (CC BY-NC 4.0)Copyright remains with the author(s), and articles are published under Creative Commons Attribution-NonCommercial 4.0 International (CC BY-NC 4.0).
