Permutation-Based RFECV on Imbalanced Students’ Grade and Students’ Dropout Datasets
DOI:
https://doi.org/10.52436/1.jutif.2026.7.4.5641Keywords:
Classification, Feature Selection, Machine Learning, PermutationRFECV, Educational Data MiningAbstract
Predicting student academic performance is a critical task in EDM. To build predictive models, it is essential to identify the most informative features from LMS data. However, existing feature selection techniques exhibit significant drawbacks. Filter Methods evaluate features independently from the learning model, ignoring interactions and context. As a result, they may select redundant features and lack adaptability to specific classifiers. To enhance classification accuracy, this research proposes a permutation-based Recursive Feature Elimination with cross-validation (RFECV) feature selection method to strengthen machine learning models' student dataset classification, thereby eliminating bias from model-based feature importance calculations using the default method. The proposed method excels on the available educational datasets (Math, Por, Student Success, and Student dataset) in terms of F1-Score. Specifically, we achieved 80.59% on math dataset and 71.81% on student success dataset, both of which are multiclass datasets.
Downloads
References
R. S. J. D. Baker and Kalina Yacef, “The State of Educational Data Mining in 2009: A Review and Future Visions,” Journal of Educational Data Mining, vol. 1, no. 1, pp. 3–16, 2009, doi: https://doi.org/10.5281/zenodo.3554657.
C. Romero, S. Ventura, and E. García, “Data mining in course management systems: Moodle case study and tutorial,” Computers & Education, vol. 51, no. 1, pp. 368–384, Aug. 2008, doi: https://doi.org/10.1016/j.compedu.2007.05.016.
M. Yağcı, “Educational data mining: prediction of students’ academic performance using machine learning algorithms,” Smart Learn. Environ., vol. 9, no. 1, p. 11, Dec. 2022, doi: 10.1186/s40561-022-00192-z.
H. Almaghrabi, B. Soh, A. Li, and I. Alsolbi, “SoK: The Impact of Educational Data Mining on Organisational Administration,” Information, vol. 15, no. 11, p. 738, Nov. 2024, doi: 10.3390/info15110738.
M. Nachouki and M. Abou Naaj, “Predicting Student Performance to Improve Academic Advising Using the Random Forest Algorithm:,” International Journal of Distance Education Technologies, vol. 20, no. 1, pp. 1–17, Mar. 2022, doi: 10.4018/IJDET.296702.
M. Hlosta, C. Herodotou, T. Papathoma, A. Gillespie, and P. Bergamin, “Predictive learning analytics in online education: A deeper understanding through explaining algorithmic errors,” Computers and Education: Artificial Intelligence, vol. 3, p. 100108, 2022, doi: 10.1016/j.caeai.2022.100108.
A. Al-Zawqari, D. Peumans, and G. Vandersteen, “A flexible feature selection approach for predicting students’ academic performance in online courses,” Computers and Education: Artificial Intelligence, vol. 3, p. 100103, 2022, doi: 10.1016/j.caeai.2022.100103.
Y. Wang, X. Zhang, and Y. Cheng, “A global and local unified feature selection algorithm based on hierarchical structure constraints,” Expert Systems with Applications, vol. 282, p. 127535, 2025, doi: https://doi.org/10.1016/j.eswa.2025.127535.
I. Guyon and A. Elisseeff, “An Introduction to Variable and Feature Selection”, doi: https://doi.org/10.1162/153244303322753616.
R. Kohavi and G. H. John, “Wrappers for feature subset selection,” Artificial Intelligence, vol. 97, no. 1–2, pp. 273–324, Dec. 1997, doi: 10.1016/S0004-3702(97)00043-X.
L. Breiman, “Random Forests,” International Journal of Advanced Computer Science and Applications, 2001.
Haibo He and E. A. Garcia, “Learning from Imbalanced Data,” IEEE Trans. Knowl. Data Eng., vol. 21, no. 9, pp. 1263–1284, Sep. 2009, doi: 10.1109/TKDE.2008.239.
Y. Liu, S. Fan, S. Xu, A. Sajjanhar, S. Yeom, and Y. Wei, “Predicting Student Performance Using Clickstream Data and Machine Learning,” Education Sciences, vol. 13, no. 1, p. 17, Dec. 2022, doi: 10.3390/educsci13010017.
A. Harif and M. A. Kassimi, “Predictive Modeling of Student Performance Using RFECV-RF for Feature Selection and Machine Learning Techniques,” ijacsa, vol. 15, no. 7, 2024, doi: 10.14569/IJACSA.2024.0150723.
M. A. Tariq, “A Study on Comparative Analysis of Feature Selection Algorithms for Students Grades Prediction,” JIOS, vol. 48, no. 1, pp. 133–147, Jun. 2024, doi: 10.31341/jios.48.1.7.
Y. Yamasari, A. Qoiriah, N. Rochmawati, K. Yoshimoto, R. A. Ahmad, and O. V. Putra, “Detecting Students’ Behavior on the E-Learning System Using SVM Kernels - Based Ensemble Learning Algorithm,” International Journal of Intelligent Engineering and Systems, vol. 16, no. 1, pp. 142–153, 2023, doi: 10.22266/ijies2023.0228.13.
S. A. Priyambada, T. Usagawa, and M. ER, “Two-layer ensemble prediction of students’ performance using learning behavior and domain knowledge,” Computers and Education: Artificial Intelligence, vol. 5, no. June, p. 100149, 2023, doi: 10.1016/j.caeai.2023.100149.
M. A. H. Alias, N. Hambali, M. A. Abdul Aziz, M. N. Taib, and R. Jailani, “Feature selection techniques and classification algorithms for student performance classification: a review,” IJECE, vol. 14, no. 3, p. 3230, Jun. 2024, doi: 10.11591/ijece.v14i3.pp3230-3243.
A. Laakel Hemdanou, M. Lamarti Sefian, Y. Achtoun, and I. Tahiri, “Comparative analysis of feature selection and extraction methods for student performance prediction across different machine learning models,” Computers and Education: Artificial Intelligence, vol. 7, p. 100301, Dec. 2024, doi: 10.1016/j.caeai.2024.100301.
P. Cortez and A. Silva, “Using Data Mining To Predict Secondary School Student Performance,” presented at the 5th Annual Future Business Technology Conference 2008, Porto: EUROSIS-ETI, Apr. 2008, pp. 5–12.
A. Altmann, L. Toloşi, O. Sander, and T. Lengauer, “Permutation importance: a corrected feature importance measure,” Bioinformatics, vol. 26, no. 10, pp. 1340–1347, May 2010, doi: 10.1093/bioinformatics/btq134.
V. Realinho, J. Machado, L. Baptista, and M. V. Martins, “Predict students’ dropout and academic success,” Dec. 2021, doi: 10.5281/zenodo.5777340.
J. Asiya, “Student Performance Prediction.” [Online]. Available: https://www.kaggle.com/datasets/asiyajan001/student-performance-prediction
C. Strobl, A.-L. Boulesteix, A. Zeileis, and T. Hothorn, “Bias in random forest variable importance measures: Illustrations, sources and a solution,” BMC Bioinformatics, vol. 8, no. 1, p. 25, Dec. 2007, doi: 10.1186/1471-2105-8-25.
C. Molnar, G. König, B. Bischl, and G. Casalicchio, “Model-agnostic feature importance and effects with dependent features: a conditional subgroup approach,” Data Min Knowl Disc, vol. 38, no. 5, pp. 2903–2941, Sep. 2024, doi: 10.1007/s10618-022-00901-9.
A. Chamma, D. A. Engemann, and B. Thirion, “Statistically Valid Variable Importance Assessment through Conditional Permutations,” Oct. 25, 2023, arXiv: arXiv:2309.07593. doi: 10.48550/arXiv.2309.07593.
Additional Files
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Irfan Pratama, Albert Yakobus Chandra, Putry Wahyu Setyaningsih, Husna Sarirah Husin, Ody Nurdiawan

This work is licensed under a Creative Commons Attribution 4.0 International License.

</a



