Building a Trustworthy Credit Default Classification Model using Interpretable Machine Learning
DOI:
https://doi.org/10.52436/1.jutif.2026.7.4.5149Keywords:
Default Prediction, Interpretable Machine Learning, SHAP, LIMEAbstract
In peer-to-peer lending, a single misjudged borrower can result in significant investor losses and erode the platform's credibility. Although machine learning models have demonstrated superior predictive performance in identifying potential loan defaults, their opaque decision-making process creates mistrust and regulatory friction, impeding their trustworthy adoption. This study addresses the critical trade-off between accuracy and interpretability by proposing an interpretable machine learning framework for trustworthy credit default classification using Bondora peer-to-peer lending data. We deploy and benchmark three classifiers —Extreme Gradient Boosting (XGBoost), Light Gradient Boosting Machine (LightGBM), and Random Forest (RF) — to assess their predictive effectiveness. To overcome the black-box nature of these models, we integrate two complementary interpretability techniques: SHapley Additive exPlanations (SHAP) for global feature attribution and Local Interpretable Model-agnostic Explanations (LIME) for granular, instance-level decision insights. The high AUC, which exceeds 0.96, and F1 score (0.937-0.940) show that all models can effectively support risk management. Our results also demonstrate that pairing high-performing classifiers with interpretability tools enhances model transparency, fosters accountability, and promotes the responsible use of machine learning in credit risk assessment.
Downloads
References
E. Katsamakas and J. M. Sánchez-Cartas, “Network Formation and Financial Inclusion in P2P Lending: A Computational Model,” Systems, vol. 10, no. 5, p. 155, Sep. 2022, doi: 10.3390/systems10050155.
I. R. Berrada, F. Barramou, and O. B. Alami, “Towards a Machine Learning-based Model for Corporate Loan Default Prediction,” Int. J. Adv. Comput. Sci. Appl., vol. 15, no. 3, 2024, doi: 10.14569/IJACSA.2024.0150357.
X. Zhu, Q. Chu, X. Song, P. Hu, and L. Peng, “Explainable prediction of loan default based on machine learning models,” Data Sci. Manag., vol. 6, no. 3, pp. 123–133, Sep. 2023, doi: 10.1016/j.dsm.2023.04.003.
S. Kumar et al., “Exploitation of Machine Learning Algorithms for Detecting Financial Crimes Based on Customers’ Behavior,” Sustainability, vol. 14, no. 21, p. 13875, Oct. 2022, doi: 10.3390/su142113875.
N. A. De Oliveira and L. F. C. Basso, “Advancing Credit Rating Prediction: The Role of Machine Learning in Corporate Credit Rating Assessment,” Risks, vol. 13, no. 6, p. 116, Jun. 2025, doi: 10.3390/risks13060116.
G. Ke et al., “LightGBM: A Highly Efficient Gradient Boosting Decision Tree,” in Proceedings of the 31st Conference on Neural Information Processing Systems, Dec. 2017.
C. Bentéjac, A. Csörgő, and G. Martínez-Muñoz, “A Comparative Analysis of XGBoost,” 2019, doi: 10.48550/ARXIV.1911.01914.
H. A. Salman, A. Kalakech, and A. Steiti, “Random Forest Algorithm Overview,” Babylon. J. Mach. Learn., vol. 2024, pp. 69–79, Jun. 2024, doi: 10.58496/BJML/2024/007.
N. Boyko and Y. Mokryk, “Detecting Fraud in Banking Transactions with Random Forest Models,” in 2021 IEEE 8th International Conference on Problems of Infocommunications, Science and Technology (PIC S&T), Kharkiv, Ukraine: IEEE, Oct. 2021, pp. 1–6. doi: 10.1109/PICST54195.2021.9772209.
A. K. Sharma, L.-H. Li, and R. Ahmad, “Default Risk Prediction Using Random Forest and XGBoosting Classifier,” in 2021 International Conference on Security and Information Technologies with AI, Internet Computing and Big-data Applications, vol. 314, G. A. Tsihrintzis, S.-J. Wang, and I.-C. Lin, Eds., in Smart Innovation, Systems and Technologies, vol. 314. , Cham: Springer International Publishing, 2023, pp. 91–101. doi: 10.1007/978-3-031-05491-4_10.
M. Karasinski and K. Bez Birolo Candiotto, “AI’s black box and the supremacy of standards,” Filos. Unisinos, vol. 25, no. 1, pp. 1–13, Mar. 2024, doi: 10.4013/fsu.2024.251.13.
H. H. Nakashima, D. Mantovani, and C. Machado Junior, “Users’ trust in black-box machine learning algorithms,” Rev. Gest., vol. 31, no. 2, pp. 237–250, Jun. 2024, doi: 10.1108/REGE-06-2022-0100.
P. Weber, K. V. Carl, and O. Hinz, “Applications of Explainable Artificial Intelligence in Finance—a systematic review of Finance, Information Systems, and Computer Science literature,” Manag. Rev. Q., vol. 74, no. 2, pp. 867–907, Jun. 2024, doi: 10.1007/s11301-023-00320-0.
P. E. De Lange, B. Melsom, C. B. Vennerød, and S. Westgaard, “Explainable AI for Credit Assessment in Banks,” J. Risk Financ. Manag., vol. 15, no. 12, p. 556, Nov. 2022, doi: 10.3390/jrfm15120556.
M. Saarela and V. Podgorelec, “Recent Applications of Explainable AI (XAI): A Systematic Literature Review,” Appl. Sci., vol. 14, no. 19, Art. no. 19, Oct. 2024, doi: 10.3390/app14198884.
A. Adadi and M. Berrada, “Peeking Inside the Black-Box: A Survey on Explainable Artificial Intelligence (XAI),” IEEE Access, vol. 6, pp. 52138–52160, 2018, doi: 10.1109/ACCESS.2018.2870052.
W. Yang et al., “Survey on Explainable AI: From Approaches, Limitations and Applications Aspects,” Hum.-Centric Intell. Syst., vol. 3, no. 3, Art. no. 3, Aug. 2023, doi: 10.1007/s44230-023-00038-y.
A. Abusitta, M. Q. Li, and B. C. M. Fung, “Survey on Explainable AI: Techniques, challenges and open issues,” Expert Syst. Appl., vol. 255, p. 124710, Dec. 2024, doi: 10.1016/j.eswa.2024.124710.
M. Bugaj, K. Wrobel, and J. Iwaniec, “Model Explainability using SHAP Values for LightGBM Predictions,” in 2021 IEEE XVIIth International Conference on the Perspective Technologies and Methods in MEMS Design (MEMSTECH), Polyana (Zakarpattya), Ukraine: IEEE, May 2021, pp. 102–106. doi: 10.1109/MEMSTECH53091.2021.9468078.
Y. Choi and E. Cha, “LightGBM Scorecard Based on SHAP Values,” SSRN Electron. J., 2023, doi: 10.2139/ssrn.4637305.
L.-H. Li, A. K. Sharma, and S.-T. Cheng, “Explainable AI based LightGBM prediction model to predict default borrower in social lending platform,” Intell. Syst. Appl., vol. 26, p. 200514, Jun. 2025, doi: 10.1016/j.iswa.2025.200514.
H. Li and W. Wu, “Loan default predictability with explainable machine learning,” Finance Res. Lett., vol. 60, p. 104867, Feb. 2024, doi: 10.1016/j.frl.2023.104867.
G. Babaei, P. Giudici, and E. Raffinetti, “Explainable FinTech lending,” J. Econ. Bus., vol. 125–126, p. 106126, May 2023, doi: 10.1016/j.jeconbus.2023.106126.
S. Yang, Z. Huang, W. Xiao, and X. Shen, “Interpretable Credit Default Prediction with Ensemble Learning and SHAP,” 2025, arXiv. doi: 10.48550/ARXIV.2505.20815.
G. F. Bone-Winkel and F. Reichenbach, “Improving credit risk assessment in P2P lending with explainable machine learning survival analysis,” Digit. Finance, vol. 6, no. 3, pp. 501–542, Sep. 2024, doi: 10.1007/s42521-024-00114-3.
P. S. Chong, J. Labadin, and F. Meziane, “Credit Risk Prediction for Peer-To-Peer Lending Platforms: An Explainable Machine Learning Approach,” J. Comput. Soc. Inform., vol. 1, no. 2, pp. 1–16, Sep. 2022, doi: 10.33736/jcsi.4761.2022.
N. Bussmann, P. Giudici, D. Marinelli, and J. Papenbrock, “Explainable AI in Fintech Risk Management,” Front. Artif. Intell., vol. 3, p. 26, Apr. 2020, doi: 10.3389/frai.2020.00026.
M. Siddhartha, “Bondora Peer-to-Peer Lending Data.” IEEE DataPort, Nov. 06, 2020. doi: 10.21227/33KZ-0S65.
Š. Lyócsa, P. Vašaničová, B. Hadji Misheva, and M. D. Vateha, “Default or profit scoring credit systems? Evidence from European and US peer-to-peer lending markets,” Financ. Innov., vol. 8, no. 1, p. 32, Dec. 2022, doi: 10.1186/s40854-022-00338-5.
S. Mondal, S. K. Shah, and V. Kumbhar, “Predicting Credit Risk in European P2P Lending: A Case Study of ‘Bondora’ Using Supervised Machine Learning Techniques,” in 2023 4th IEEE Global Conference for Advancement in Technology (GCAT), Bangalore, India: IEEE, Oct. 2023, pp. 1–6. doi: 10.1109/GCAT59970.2023.10353353.
Y. Liu, L. J. Baals, J. Osterrieder, and B. Hadji-Misheva, “Leveraging network topology for credit risk assessment in P2P lending: A comparative study under the lens of machine learning,” Expert Syst. Appl., vol. 252, p. 124100, Oct. 2024, doi: 10.1016/j.eswa.2024.124100.
M. Madaan, A. Kumar, C. Keshri, R. Jain, and P. Nagrath, “Loan default prediction using decision trees and random forest: A comparative study,” IOP Conf. Ser. Mater. Sci. Eng., vol. 1022, no. 1, p. 012042, Jan. 2021, doi: 10.1088/1757-899X/1022/1/012042.
R. Kurniawan, “Application of Random Forest Algorithm on Credit Risk Analysis,” Procedia Comput. Sci., vol. 245, pp. 740–749, 2024, doi: 10.1016/j.procs.2024.10.300.
S. Hartini, Z. Rustam, G. S. Saragih, and M. J. Segovia Vargas, “Estimating probability of banking crises using random forest,” IAES Int. J. Artif. Intell. IJ-AI, vol. 10, no. 2, p. 407, Jun. 2021, doi: 10.11591/ijai.v10.i2.pp407-413.
J. Wang, W. Rong, Z. Zhang, and D. Mei, “Credit Debt Default Risk Assessment Based on the XGBoost Algorithm: An Empirical Study from China,” Wirel. Commun. Mob. Comput., vol. 2022, pp. 1–14, Mar. 2022, doi: 10.1155/2022/8005493.
D. Wang, L. Li, and D. Zhao, “Corporate finance risk prediction based on LightGBM,” Inf. Sci., vol. 602, pp. 259–268, Jul. 2022, doi: 10.1016/j.ins.2022.04.058.
L. Sari, A. Romadloni, R. Lityaningrum, and H. D. Hastuti, “Implementation of LightGBM and Random Forest in Potential Customer Classification,” TIERS Inf. Technol. J., vol. 4, no. 1, pp. 43–55, Jun. 2023, doi: 10.38043/tiers.v4i1.4355.
A. D. Hartanto, Y. Nur Kholik, and Y. Pristyanto, “Stock Price Time Series Data Forecasting Using the Light Gradient Boosting Machine (LightGBM) Model,” JOIV Int. J. Inform. Vis., vol. 7, no. 4, p. 2270, Dec. 2023, doi: 10.62527/joiv.7.4.1740.
Ş. K. Çorbacıoğlu and G. Aksel, “Receiver operating characteristic curve analysis in diagnostic accuracy studies: A guide to interpreting the area under the curve value,” Turk. J. Emerg. Med., vol. 23, no. 4, pp. 195–198, Oct. 2023, doi: 10.4103/tjem.tjem_182_23.
S. Lundberg and S.-I. Lee, “A Unified Approach to Interpreting Model Predictions,” 2017, arXiv. doi: 10.48550/ARXIV.1705.07874.
M. T. Ribeiro, S. Singh, and C. Guestrin, “‘Why Should I Trust You?’: Explaining the Predictions of Any Classifier,” 2016, arXiv. doi: 10.48550/ARXIV.1602.04938.
Additional Files
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Haikal Djauhari, Anton Yudhana, Sunardi

This work is licensed under a Creative Commons Attribution 4.0 International License.

</a



