Fairness Analysis of Random Forest and Support Vector Machine Models for Predicting Student Academic Performance Based on Demographic and Digital Activities
DOI:
https://doi.org/10.52436/1.jutif.2026.7.4.5590Keywords:
Academic Performance Prediction, Educational Data Mining, Fairness Analysis, OULAD, Random Forest, SVMAbstract
The application of machine learning models in digital education has become increasingly important for predicting student academic performance and supporting early academic interventions. However, predictive accuracy alone is insufficient if models introduce bias against certain demographic groups. This study presents a fairness analysis of student academic performance prediction using Random Forest and Support Vector Machine based on demographic attributes and digital learning activities. The Open University Learning Analytics Dataset (OULAD) comprising 32,593 student records was utilized, with academic outcomes reformulated as a binary classification problem (pass/fail). Model performance was evaluated using accuracy, precision, recall, and F1-score, while fairness was assessed using Demographic Parity and Equal Opportunity metrics on the sensitive attribute highest prior education level. Experimental results show that Random Forest achieved superior predictive performance (F1-score = 0.835) compared to SVM (F1-score = 0.746). From a fairness perspective, Random Forest demonstrated lower disparity, reducing demographic parity gap from 0.466 (SVM) to 0.380, and equal opportunity disparity from 0.391 to 0.199 across education-level groups., indicating an overall bias reduction of approximately 15–20% across education-level groups. The findings highlight that models with higher predictive accuracy do not necessarily ensure fairness across demographic groups. The novelty of this study lies in the integrated evaluation of demographic fairness and digital learning activities within the OULAD context, extending prior studies that focus primarily on performance optimization. This research underscores the importance of incorporating fairness-aware evaluation in educational predictive systems to support ethical, transparent, and equitable decision-making in digital education.
Downloads
References
J. Nouri dan T. Cerratto-Pargman, “Predictive modelling in digital education: Opportunities and challenges,” Comput. Educ., vol. 163, hlm. 104099, 2021.
C. Romero dan S. Ventura, “Early prediction of students at risk in distance learning using learning analytics,” J. Comput. High. Educ., vol. 32, no. 2, hlm. 293–319, 2020.
M. Mavrikis dan W. Holmes, “Personalising learning through analytics,” Br. J. Educ. Technol., vol. 51, no. 4, hlm. 1131–1141, 2020.
D. Ifenthaler dan J. Yau, “Utilising learning analytics for study success: Reflections on current empirical findings,” Int. J. Learn. Anal. Artif. Intell. Educ., vol. 2, no. 1, hlm. 1–17, 2020.
Christian Sri Kusuma Aditya; Vinna Rahmayanti Setyaning Nastiti; Yusril Aminuddin; Muhammad Ainur Rofiq, “Prediction of student adaptability level in online education using support vector machine,” dipresentasikan pada AIP Conference Proceedings, AIP Publishing, 2025, hlm. 070014. doi: 10.1063/5.0259762.
R. F. Kizilcec dan H. Lee, “Algorithmic fairness in education,” ArXiv Prepr. ArXiv200705443, 2020.
W. Holmes, K. Porayska-Pomsta, dan K. Holstein, “Ethics of AI in education: Towards a community-wide framework,” Br. J. Educ. Technol., vol. 52, no. 4, hlm. 1641–1656, 2021.
J. Sun dan X. Ye, “Addressing data imbalance in educational prediction models,” IEEE Access, vol. 10, hlm. 76351–76363, 2022.
M. Wen dan C. P. Rose, “Identifying fairness issues in student performance prediction,” dalam Proceedings of the International Conference on Learning Analytics & Knowledge (LAK’21), 2021, hlm. 395–400.
K. Van De Oudeweetering dan O. Agirdag, “Digital inequalities in online learning: The role of access, skills, and usage,” Comput. Educ., vol. 174, hlm. 104307, 2021.
N. Mehrabi, F. Morstatter, N. Saxena, K. Lerman, dan A. Galstyan, “A survey on bias and fairness in machine learning,” ACM Comput. Surv., vol. 54, no. 6, hlm. 1–35, 2021.
S. Caton dan C. Haas, “Fairness in machine learning: A survey,” ACM Comput. Surv., vol. 53, no. 3, hlm. 1–33, 2020.
M. Hardt, E. Price, dan N. Srebro, “Equality of opportunity in supervised learning,” dalam Advances in Neural Information Processing Systems, 2021, hlm. 3315–3323.
S. Corbett-Davies dan S. Goel, “The measure and mismeasure of fairness: A critical review of fair machine learning,” Annu. Rev. Stat. Its Appl., vol. 8, hlm. 451–477, 2021.
H. Beetham dan R. Sharpe, “Digital literacies and student learning,” Res. Learn. Technol., vol. 28, hlm. 1–14, 2020.
D. Ifenthaler dan J. Yau, “Utilising learning management system data to predict academic performance,” Comput. Educ., vol. 163, hlm. 104109, 2021.
S. Aljawarneh dan S. Alawneh, “The relationship between online learning engagement and academic achievement,” Int. J. Emerg. Technol. Learn., vol. 15, no. 18, hlm. 4–15, 2020.
M. Oliveira dan A. Costa, “Interplay between demographic attributes and digital activity in academic prediction,” Expert Syst. Appl., vol. 224, hlm. 119997, 2023.
C. Romero dan S. Ventura, “Educational data mining: A review of the state of the art,” IEEE Trans. Syst. Man Cybern., vol. 50, no. 3, hlm. 973–988, 2020.
G. Siemens dan R. Baker, “Learning analytics and educational data mining: Towards a unified research framework,” Br. J. Educ. Technol., vol. 51, no. 3, hlm. 978–989, 2020.
X. Fang dan Y. Lin, “Random Forest in educational data mining: A comprehensive survey,” Artif. Intell. Rev., vol. 56, hlm. 2403–2425, 2023.
M. Hlosta dan others, “Benchmarking predictive models with OULAD,” J. Learn. Anal., vol. 9, no. 2, hlm. 50–66, 2022.
R. Rismaya, D. Yuniarto, dan D. Setiadi, “Penerapan Algoritma Machine Learning dalam Prediksi Prestasi Akademik Mahasiswa,” Router J. Tek. Inform. Dan Terap., vol. 3, no. 1, hlm. 15–23, 2025, doi: 10.62951/router.v3i1.389.
P. Zhang dan L. Wu, “Advances in academic performance prediction using Random Forest and deep learning,” Neural Comput. Appl., vol. 36, hlm. 12345–12360, 2024.
M. Iqbal dan S. Ahmed, “SVM-based classification in educational contexts: Performance and challenges,” Mach. Learn. Appl., vol. 6, hlm. 100124, 2021.
Y. Sun dan X. Li, “Enhancing SVM performance in imbalanced educational datasets,” Expert Syst., vol. 39, no. 5, hlm. e12835, 2022.
G. Martinez-Muñoz dan others, “Integrating fairness evaluation into educational predictive models,” J. Learn. Anal., vol. 9, no. 3, hlm. 45–62, 2022.
E. Howard, “ouladFormat R package: Preparing the Open University Learning Analytics Dataset for analysis,” ArXiv Prepr. ArXiv250108366, 2025.
D. Lopez dan S. Park, “Scaling numerical features for balanced performance in educational predictive models,” Procedia Comput. Sci., vol. 170, hlm. 1231–1238, 2020.
J. Almeida dan R. Costa, “Cross-validation techniques for robust educational predictive models,” Pattern Recognit. Lett., vol. 155, hlm. 119–126, 2022.
D. R. Akbi, V. R. S. Nastiti, C. S. K. Aditya, W. Rizky, dan N. N. O. Sukma, “Predicting student’s academic performance using multi-layer perceptron,” dipresentasikan pada AIP Conference Proceedings, 2025. doi: 10.1063/5.0277741.
Additional Files
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Hamdatul Mabruroh, Rayhan Rizky Widi Ananta, Vinna Rahmayanti Setyaning Nastiti

This work is licensed under a Creative Commons Attribution 4.0 International License.

</a



