Explainable Machine Learning of Scam Solicitation Exposure Among Indonesian Mobile Phone Users Using Global Findex 2025

Authors

  • Rahmadani Universitas Muslim Indonesia Author
  • Lukman Syafie Universiti Kuala Lumpur Author

DOI:

https://doi.org/10.56705/0kk16335

Keywords:

Digital Scam, Scam Solicitation, Global Findex 2025, Explainable Machine Learning, Random Forest, SHAP, Digital Financial Inclusion, Indonesia

Abstract

Introduction: The expansion of mobile connectivity and digital financial services has increased financial inclusion while creating additional channels for scam solicitation. This study examines factors associated with reported scam-solicitation exposure among Indonesian mobile phone users using nationally representative Global Findex 2025 data. Method: The analysis included 850 respondents with valid scam-solicitation responses, consisting of 160 exposed and 690 non-exposed individuals. Survey weights were incorporated into descriptive analysis, model fitting, and evaluation. Logistic Regression, Random Forest, and XGBoost were assessed using stratified five-fold cross-validation. Four incremental predictor sets representing demographics, connectivity and digital capability, digital behavior, and financial participation were evaluated. SHAP was used for model interpretation. Results: The survey-weighted prevalence of reported scam-solicitation exposure was 18.54%. Random Forest achieved the strongest overall performance, with a ROC-AUC of 0.677, PR-AUC of 0.331, sensitivity of 0.590, and Brier score of 0.143. The demographic-only model achieved a ROC-AUC of 0.539, increasing to 0.617 with connectivity variables, 0.650 with digital behavior, and 0.677 with financial-participation variables. SHAP identified online activities, mobile-money ownership, digitally enabled accounts, digital payments, and online job searching among the most influential predictors. Conclusion: Digital participation provides substantially greater discriminatory information about scam-solicitation exposure than demographic characteristics alone. These findings support digital-safety interventions embedded within high-engagement digital and financial environments.

References

[1] L. Klapper, D. Singer, L. Starita, and A. Norris, The Global Findex Database 2025: Connectivity and Financial Inclusion in the Digital Economy. Washington, DC, USA: World Bank, 2025, doi: https://doi.org/10.1596/978-1-4648-2204-9.

[2] P. Prasetyoputra, Y. Isnasari, A. P. S. Prasojo, and I. Hermawan, “Factors associated with financial inclusion in Indonesia before and during COVID-19: Evidence from Global Findex data,” Statistics, Optimization & Information Computing, vol. 15, no. 1, pp. 93–127, 2026, doi: https://doi.org/10.19139/soic-2310-5070-2852.

[3] M. Houtti, A. Roy, V. N. R. Gangula, and A. M. Walker, “A survey of scam exposure, victimization, types, vectors, and reporting in 12 countries,” Journal of Online Trust and Safety, vol. 2, no. 4, 2024, doi: https://doi.org/10.54501/jots.v2i4.204.

[4] G. Norris, A. Brookes, and D. Dowell, “The psychology of internet fraud victimisation: A systematic review,” Journal of Police and Criminal Psychology, vol. 34, no. 3, pp. 231–245, 2019, doi: https://doi.org/10.1007/s11896-019-09334-5.

[5] T. C. Pratt, K. Holtfreter, and M. D. Reisig, “Routine online activity and internet fraud targeting: Extending the generality of routine activity theory,” Journal of Research in Crime and Delinquency, vol. 47, no. 3, pp. 267–296, 2010, doi: https://doi.org/10.1177/0022427810365903.

[6] M. T. Whitty, “Predicting susceptibility to cyber-fraud victimhood,” Journal of Financial Crime, vol. 26, no. 1, pp. 277–292, 2019, doi: https://doi.org/10.1108/JFC-10-2017-0095.

[7] V. Balakrishnan, U. Ahhmed, and F. Basheer, “Personal, environmental and behavioral predictors associated with online fraud victimization among adults,” PLOS ONE, vol. 20, no. 1, Art. no. e0317232, 2025, doi: https://doi.org/10.1371/journal.pone.0317232.

[8] T. Kraiwanit, P. Limna, R. Sonsuphap, and V. Chutipat, “Predictors of digital fraud: Evidence from Thailand,” Journal of Risk and Financial Management, vol. 18, no. 12, Art. no. 671, 2025, doi: https://doi.org/10.3390/jrfm18120671.

[9] A. Ardoni, “Information items used by online fraudsters and its relationship to Indonesian digital literacy,” Berkala Ilmu Perpustakaan dan Informasi, vol. 18, no. 2, pp. 326–337, 2022, doi: https://doi.org/10.22146/bip.v18i2.5866.

[10] S. Lestari, W. R. Adawiyah, A. L. Alhamidi, J. Prayogi, and R. Haryanto, “Navigating perilous seas: Unmasking online banking frauds, perceived usefulness, fear of cybercrime and distrust in online banking,” Safer Communities, vol. 23, no. 4, pp. 444–464, 2024, doi: https://doi.org/10.1108/SC-04-2024-0018.

[11] H. Simaremare, M. Fikri, and Z. Shukur, “Confirmatory factor analysis of phishing susceptibility in Indonesia,” International Journal of Advanced Computer Science and Applications, vol. 17, no. 6, 2026, doi: https://doi.org/10.14569/IJACSA.2026.0170638.

[12] A. Asran, E. I. Alwi, and H. Azis, “Implementasi metode static forensics untuk ekstraksi file steganografi pada bukti digital menggunakan framework NIST,” JIPI (Jurnal Ilmiah Penelitian dan Pembelajaran Informatika), vol. 11, no. 1, pp. 544–556, 2026, doi: https://doi.org/10.29100/jipi.v11i1.7406.

[13] M. Lokanan and S. Liu, “Predicting fraud victimization using classical machine learning,” Entropy, vol. 23, no. 3, Art. no. 300, 2021, doi: https://doi.org/10.3390/e23030300.

[14] H. Wang and D. Wang, “Constructing online fraud victimization models: A machine learning approach,” Journal of Forensic Psychology Research and Practice, early access, 2026, doi: https://doi.org/10.1080/24732850.2026.2643611.

[15] G. S. Erbuğa and C. Ünal, “Explainable artificial intelligence-driven fraud detection: Behavioral proxies of digital financial literacy in online payment models,” Borsa Istanbul Review, Art. no. 100840, 2026, doi: https://doi.org/10.1016/j.bir.2026.100840.

[16] Y. Zhou, H. Li, Z. Xiao, and J. Qiu, “A user-centered explainable artificial intelligence approach for financial fraud detection,” Finance Research Letters, vol. 58, Part A, Art. no. 104309, 2023, doi: https://doi.org/10.1016/j.frl.2023.104309.

[17] U. Zafar and F. Wu, “Methodological challenges in explainable AI for fraud detection: A systematic literature review,” Artificial Intelligence Review, vol. 59, Art. no. 115, 2026, doi: https://doi.org/10.1007/s10462-026-11516-7.

[18] H. Azis, Y. Salim, and Y. Puspitasari, “Deep learning approaches for soft tissue classification of cephalometric images in orthodontics,” JOIV: International Journal on Informatics Visualization, vol. 10, no. 3, pp. 988–995, 2026, doi: https://doi.org/10.62527/joiv.10.3.3885.

[19] H. Azis, R. A. Jalil, and A. R. Manga’, “Comparative performance of VGG16 and EfficientNetB0-based transfer learning for brain tumor classification,” Knowledge Engineering and Data Science, vol. 8, no. 2, pp. 185–197, 2025, doi: https://doi.org/10.17977/um018v8i22025p185-197.

[20] A. R. Manga, A. P. Utami, H. Azis, Y. Salim, and A. Faradibah, “Optimizing classification models for medical image diagnosis: A comparative analysis on multi-class datasets,” Computer Science and Information Technologies, vol. 5, no. 3, pp. 205–214, 2024, doi: https://doi.org/10.11591/csit.v5i3.p205-214.

[21] H. Azis and N. Rismayanti, “Prediksi anemia dari pixel gambar dan level hemoglobin menggunakan Random Forest Classifier,” NERO (Networking Engineering Research Operation), vol. 9, no. 1, pp. 21–34, 2024, doi: https://doi.org/10.21107/nero.v9i1.27916.

[22] Aldo, A. W. Murdiyanto, and U. S. Aesyi, “Zero-shot detection of IndoT5-synthesized Indonesian scientific abstracts using mDeBERTa v3,” Indonesian Journal of Data and Science, vol. 7, no. 2, pp. 333–349, 2026, doi: https://doi.org/10.56705/ijodas.v7i2.457.

[23] E. H. Saputra, I. A. E. Zaeni, D. D. Prasetya, A. M. Zain, W. Antonius, and I. M. Wirawan, “Comparative evaluation of machine learning models for heavy crude oil viscosity prediction using repeated nested cross-validation and independent holdout testing,” Indonesian Journal of Data and Science, vol. 7, no. 2, 2026, doi: https://doi.org/10.56705/ijodas.v7i2.455.

[24] M. R. Hardika and B. Miftahurrohmah, “Weakly supervised sentiment analysis of gold price discussions using conventional machine learning and IndoBERT,” Indonesian Journal of Data and Science, vol. 7, no. 2, pp. 274–290, 2026, doi: https://doi.org/10.56705/ijodas.v7i2.445.

[25] D. H. Zulfikar, E. S. J. Atmadji, and B. S. W. Poetro, “Leakage-aware and explainable machine learning for healthcare claim fraud detection using imbalanced medical insurance data,” International Journal of Artificial Intelligence in Medical Issues, vol. 4, no. 1, pp. 19–34, 2026, doi: https://doi.org/10.56705/z3207345.

[26] Development Research Group, Finance and Private Sector Development Unit, Indonesia—The Global Findex Database 2025: Connectivity and Financial Inclusion in the Digital Economy, Ref. IDN_2024_FINDEX_v02_M, World Bank, 2025, doi: https://doi.org/10.48529/CDK5-2M94.

[27] L. Breiman, “Random forests,” Machine Learning, vol. 45, pp. 5–32, 2001, doi: https://doi.org/10.1023/A:1010933404324.

[28] T. Chen and C. Guestrin, “XGBoost: A scalable tree boosting system,” in Proc. 22nd ACM SIGKDD Int. Conf. Knowledge Discovery and Data Mining (KDD ’16), San Francisco, CA, USA, 2016, pp. 785–794, doi: https://doi.org/10.1145/2939672.2939785.

[29] W. J. Youden, “Index for rating diagnostic tests,” Cancer, vol. 3, no. 1, pp. 32–35, 1950, doi: https://doi.org/10.1002/1097-0142(1950)3:1<32::AID-CNCR2820030106>3.0.CO;2-3.

[30] G. W. Brier, “Verification of forecasts expressed in terms of probability,” Monthly Weather Review, vol. 78, no. 1, pp. 1–3, 1950, doi: https://doi.org/10.1175/1520-0493(1950)078<0001>2.0.CO;2.

[31] T. Saito and M. Rehmsmeier, “The precision-recall plot is more informative than the ROC plot when evaluating binary classifiers on imbalanced datasets,” PLOS ONE, vol. 10, no. 3, Art. no. e0118432, 2015, doi: https://doi.org/10.1371/journal.pone.0118432.

[32] S. M. Lundberg and S.-I. Lee, “A unified approach to interpreting model predictions,” in Advances in Neural Information Processing Systems 30, 2017, pp. 4765–4774.

[33] A. Halid, D. A. Purnamasari, A. C. Saputra, and N. M. Setiohardjo, “Explainable machine learning for predicting the mental health impact of AI and digital platform usage among students,” International Journal of Artificial Intelligence in Medical Issues, vol. 4, no. 1, pp. 1–18, 2026, doi: https://doi.org/10.56705/pxn6qg39.

Downloads

Published

2026-04-30