Comparative Evaluation of Machine Learning and Ensemble Learning with Synthetic Minority Oversampling and Random Undersampling for Loan Approval Prediction Across Multiple Datasets
DOI:
https://doi.org/10.52436/1.jutif.2026.7.4.5854Keywords:
class imbalance, ensemble learning, loan approval prediction, random undersampling, synthetic minority oversamplingAbstract
Loan approval prediction remains a critical yet challenging task in financial risk management, particularly due to class imbalance and variability across lending datasets. This study presents a comparative evaluation of eleven machine learning and ensemble learning algorithms for loan approval prediction, evaluated across three publicly available datasets with distinct characteristics to ensure generalizability of findings. The algorithms evaluated include Logistic Regression, Decision Tree, Random Forest, Soft Voting, Hard Voting, AdaBoost, Gradient Boosting, XGBoost, Bagging, Pasting, and Stacking. To address class imbalance, three data balancing strategies are applied: original imbalanced data, Synthetic Minority Oversampling Technique, and random undersampling. Model performance is assessed using accuracy, precision, recall, and F1-score under 5-fold cross-validation with a 75:25 stratified train-test split. Experimental results demonstrate that ensemble methods consistently outperform single classifiers across all datasets. Stacking achieves the highest accuracy on original data across all three datasets (88.33%, 76.00%, and 93.16%, respectively), while XGBoost demonstrates competitive and stable performance across multiple data balancing conditions. Results also reveal that Synthetic Minority Oversampling does not universally improve model performance, as certain dataset characteristics lead to accuracy degradation under oversampling conditions. These findings contribute to the field of computer science by establishing a multi-dataset benchmark for ensemble learning in automated credit scoring and by providing empirical evidence on the interaction between data balancing strategies and classifier architectures, offering a practical framework for selecting robust machine learning pipelines in financial decision support systems.
Downloads
References
Q. Zhang, "Loan risk prediction model based on random forest," Advances in Economics, Management and Political Sciences, vol. 5, pp. 216–222, 2023, doi: 10.54254/2754-1169/5/20220082.
Y. Guan, Y. Wang, and Y. Yang, "Personal loan default prediction analysis based on TabNet and logistic regression," Applied and Computational Engineering, vol. 103, pp. 187–195, 2024, doi: 10.54254/2755-2721/103/20241160.
M. R. Islam, M. S. Islam, S. Shama, A. I. Chowdhury, and M. M. H. Lamyea, "Enhancing bank loan approval efficiency using machine learning: An ensemble model approach," Engineering and Technology Journal, vol. 9, no. 7, 2024, doi: 10.47191/etj/v9i07.24.
R. Yang, "Machine learning-based loan default prediction in peer-to-peer lending," Highlights in Science, Engineering and Technology, vol. 94, pp. 310–318, 2024, doi: 10.54097/qdjd8r65.
N. Uddin et al., "An ensemble machine learning based bank loan approval predictions system with a smart application," International Journal of Cognitive Computing in Engineering, vol. 4, pp. 327–339, 2023, doi: 10.1016/j.ijcce.2023.09.001.
S. Ramadhani and M. Wayahdi, "K-nearest neighbor and random forest algorithms in loan approval prediction," Jurnal Minfo Polgan, vol. 13, pp. 1307–1313, Dec. 2024, doi: 10.33395/jmp.v13i1.14345.
F. Lwesya and A. B. S. Mwakalobo, "Frontiers in microfinance research for small and medium enterprises (SMEs) and microfinance institutions (MFIs): A bibliometric analysis," Future Business Journal, vol. 9, no. 1, p. 17, 2023, doi: 10.1186/s43093-023-00195-3.
A. A. Egwa, B. Habeeb, A. A. Ahmad, and S. M. Bizi, "Default prediction for loan lenders using machine learning algorithms," Sule Lamido University Journal of Science and Technology, vol. 5, no. 1–2, pp. 1–12, Dec. 2022, doi: 10.56471/slujst.v5i.222.
D. Dansana, S. G. K. Patro, B. K. Mishra, V. Prasad, A. Razak, and A. W. Wodajo, "Analyzing the impact of loan features on bank loan prediction using random forest algorithm," Engineering Reports, vol. 6, no. 2, p. e12707, 2024, doi: 10.1002/eng2.12707.
B. Nugroho, Purwanto, and H. Himawan, "Enhancing default prediction in P2P lending using random forest and Grey Wolf Optimization-based feature selection," Journal of Applied Intelligent System, vol. 8, pp. 363–375, Nov. 2023, doi: 10.33633/jais.v8i3.9234.
J. W. Lee and S. Y. Sohn, "Evaluating borrowers' default risk with a spatial probit model reflecting the distance in their relational network," PLOS ONE, vol. 16, no. 12, p. e0261737, 2021, doi: 10.1371/journal.pone.0261737.
L. Nguyen, M. Ahsan, and J. Haider, "Reimagining peer-to-peer lending sustainability: Unveiling predictive insights with innovative machine learning approaches for loan default anticipation," FinTech, vol. 3, no. 1, pp. 184–215, 2024, doi: 10.3390/fintech3010012.
D. Maloney, S.-C. Hong, and B. N. Nag, "Economic disruptions in repayment of peer loans," International Journal of Financial Studies, vol. 11, no. 4, p. 116, 2023, doi: 10.3390/ijfs11040116.
S. Mor, R. Aneja, S. Madan, and S. Gupta, "Artificial intelligence and loan default: The case of commercial banks in India," Strategic Change, vol. 31, no. 6, pp. 571–580, 2022, doi: 10.1002/jsc.2529.
K. Kohv and O. Lukason, "What best predicts corporate bank loan defaults? An analysis of three different variable domains," Risks, vol. 9, no. 2, p. 29, 2021, doi: 10.3390/risks9020029.
N. Alagic et al., "Machine learning for an enhanced credit risk analysis: A comparative study of loan approval prediction models integrating mental health data," Machine Learning and Knowledge Extraction, vol. 6, no. 1, pp. 53–77, 2024, doi: 10.3390/make6010004.
S. Bhatore, L. Mohan, and Y. R. Reddy, "Machine learning techniques for credit risk evaluation: A systematic literature review," Journal of Banking and Financial Technology, vol. 4, no. 1, pp. 111–138, 2020, doi: 10.1007/s42786-020-00020-3.
S. Kokate and M. S. R. Chetty, "Credit risk assessment of loan defaulters in commercial banks using voting classifier ensemble learner machine learning model," International Journal of Safety and Security Engineering, vol. 11, no. 5, pp. 565–572, 2021, doi: 10.18280/ijsse.110508.
C. X. Pham, H. N. Trinh, and L. Q. Tran, "A robust approach to credit scoring with deep learning and embedded methods," Engineering, Technology & Applied Science Research, vol. 15, no. 6, pp. 29284–29291, 2025, doi: 10.48084/etasr.12649.
M. Jumaa, M. Saqib, and A. Attar, "Improving credit risk assessment through deep learning-based consumer loan default prediction model," International Journal of Finance and Banking Studies, vol. 12, no. 1, pp. 85–92, 2023, doi: 10.20525/ijfbs.v12i1.2579.
W. F. Setiawan, A. Amirullah, I. P. Ariatama, and R. N. E. Anggraini, "CNN-LSTM architecture for multi-task sentiment and emotion classification on large-scale Indonesian TikTok application reviews," CommIT (Communication and Information Technology) Journal, vol. 20, no. 1, pp. 77–91, Mar. 2026.
A. B. Dina et al., "Comparison of oversampling techniques in prediction judicial decisions of divorce trials in family courts," in Proc. 2024 Int. Conf. Information Technology Research and Innovation (ICITRI), Jakarta, Indonesia, 2024, pp. 13–18, doi: 10.1109/ICITRI62858.2024.10699016.
L. T. Trinh, "A comparative analysis of consumer credit risk models in peer-to-peer lending," Journal of Economics, Finance and Administrative Science, vol. 29, no. 58, pp. 346–365, 2024, doi: 10.1108/jefas-04-2021-0026.
R. Hlongwane, K. K. K. M. Ramaboa, and W. T. Mongwe, "Enhancing credit scoring accuracy with a comprehensive evaluation of alternative data," PLOS ONE, vol. 19, no. 5, p. e0303566, 2024, doi: 10.1371/journal.pone.0303566.
X. Dong, W. Xue, and J. Chen, "Analysis and comparison of loan default prediction models based on XGBoost and LightGBM algorithm," Academic Journal of Computing and Information Science, vol. 6, no. 9, 2023, doi: 10.25236/ajcis.2023.060905.
R. N. E. Anggraini, S. Quinevera, R. Mardianto, R. Anggoro, A. M. Shiddiqi, and A. Yuniarti, "Predictive modeling of term deposit subscriptions using machine learning and bagging ensemble methods on imbalanced data," in Proc. 2025 Int. Conf. Smart-Green Technology in Electrical and Information Systems (ICSGTEIS), Denpasar, Indonesia, 2025, pp. 295–300, doi: 10.1109/ICSGTEIS68532.2025.11284473.
R. N. E. Anggraini et al., "A decision tree knowledge-based system for reviewing research ethics protocol," in Proc. 2022 IEEE Int. Conf. Communication, Networks and Satellite (COMNETSAT), Solo, Indonesia, 2022, pp. 50–55, doi: 10.1109/COMNETSAT56033.2022.9994414.
W. F. Setiawan, A. Amirullah, I. P. Ariatama, and R. N. E. Anggraini, "Pemanfaatan pembelajaran mesin untuk klasifikasi kebutuhan perangkat lunak," Jurnal Teknologi Informasi dan Ilmu Komputer (JTIIK), vol. 13, no. 1, pp. 11–20, Feb. 2026, doi: 10.25126/jtiik.2026131.
N. V. Chawla, K. W. Bowyer, L. O. Hall, and W. P. Kegelmeyer, "SMOTE: Synthetic minority over-sampling technique," Journal of Artificial Intelligence Research, vol. 16, pp. 321–357, 2002, doi: 10.1613/jair.953.
D. H. Wolpert, "Stacked generalization," Neural Networks, vol. 5, no. 2, pp. 241–259, 1992, doi: 10.1016/S0893-6080(05)80023-1.
T. Chen and C. Guestrin, "XGBoost: A scalable tree boosting system," in Proc. 22nd ACM SIGKDD Int. Conf. Knowledge Discovery and Data Mining, San Francisco, CA, USA, 2016, pp. 785–794, doi: 10.1145/2939672.2939785.
H. Hofmann, "Statlog (German Credit Data)," UCI Machine Learning Repository, 1994. [Online]. Available: https://archive.ics.uci.edu/dataset/144/statlog+german+credit+data.
L. Breiman, "Random forests," Machine Learning, vol. 45, no. 1, pp. 5–32, 2001, doi: 10.1023/A:1010933404324.
Additional Files
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Ilham Putra Ariatama, Wahyu Fajar Setiawan, Afif Amirullah, Ratih Nur Esti Anggraini

This work is licensed under a Creative Commons Attribution 4.0 International License.

</a



