Perbandingan Algoritma Supervised Learning dalam Memprediksi Jalur Seleksi Masuk Mahasiswa di Universitas Negeri Gorontalo
Abstract
The choice of university admission pathway (SNBT, SNBP, or independent selection/Mandiri) is closely related to students' demographic and administrative characteristics, and understanding this pattern benefits higher education institutions in designing more targeted recruitment strategies. This study aims to compare the performance of four supervised learning algorithms Decision Tree, Random Forest, Naïve Bayes, and k-Nearest Neighbor (k-NN) in predicting the admission pathway of Informatics Engineering students at Universitas Negeri Gorontalo (UNG). Data were obtained from UNG's Integrated Academic Information System (SIAT) for 1,237 students from the Information Technology Education and Information Systems study programs (2018–2024 cohorts), which after data cleaning resulted in 1,232 samples across three pathway classes (SNBT = 682, SNBP = 417, Mandiri = 133). Feature selection using information gain identified five informative features: selection type (national/local), age, cohort year, initial registration semester, and gender. Evaluation was conducted through three scenarios: 80:20 hold-out with Synthetic Minority Oversampling Technique (SMOTE), and 10-fold stratified cross-validation, followed by a Friedman significance test. The 10-fold cross-validation results show mean accuracies of 64.61% for Random Forest, 64.37% for Decision Tree, 63.47% for Naïve Bayes, and 61.03% for k-NN. The Friedman test indicated no statistically significant difference in performance among the four algorithms (χ² = 4.39; p = 0.222). The confusion matrix revealed that all algorithms perfectly classified the Mandiri class due to a definitional dependency with the selection-type feature, while most misclassifications occurred between the SNBP and SNBT classes. Applying SMOTE consistently improved minority-class (SNBP) recall across all four algorithms, but its effect on macro F1-score and overall accuracy varied by algorithm and was not uniformly positive. This study recommends incorporating academic behavioral features to improve the model's discriminative ability between national selection pathways.
Downloads
References
Arévalo-Cordovilla, F. E., & Peña, M. (2024). Comparative Analysis of Machine Learning Models for Predicting Student Success in Online Programming Courses: A Study Based on LMS Data and External Factors. Mathematics, 12(20), 3272. https://doi.org/10.3390/math12203272
Azizah, Z., Ohyama, T., Zhao, X., Ohkawa, Y., & Mitsuishi, T. (2024). Predicting at-risk students in the early stage of a blended learning course via machine learning using limited data. Computers and Education: Artificial Intelligence, 7, 100261. https://doi.org/10.1016/j.caeai.2024.100261
Chawla, N. V., Bowyer, K. W., Hall, L. O., & Kegelmeyer, W. P. (2002). SMOTE: Synthetic Minority Over-sampling Technique. Journal of Artificial Intelligence Research, 16, 321-357. https://doi.org/10.1613/jair.953
Elbouknify, I., Berrada, I., Mekouar, L., Iraqi, Y., Bergou, E. H., Belhabib, H., Nail, Y., & Wardi, S. (2025). AI-based identification and support of at-risk students: A case study of the Moroccan education system [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2504.07160
Friedman, M. (1937). The Use of Ranks to Avoid the Assumption of Normality Implicit in the Analysis of Variance. Journal of the American Statistical Association, 32(200), 675-701. https://doi.org/10.2307/2279372
Indra, & Agustinawati, D. (2024). Perbandingan Metode Data Mining pada Prediksi Kelulusan Mahasiswa Fakultas Teknologi Informasi di Perguruan Tinggi dengan Algoritma Naive Bayes dan K-Nearest Neighbor. SIMETRIS, 15(1), 69-84.
Matz, S. C., Bukow, C. S., Peters, H., et al. (2023). Using machine learning to predict student retention from socio-demographic characteristics and app-based engagement metrics. Scientific Reports, 13, 5705. https://doi.org/10.1038/s41598-023-32484-w
Oppong, S. (2023). Predicting Students' Performance Using Machine Learning Algorithms: A Review. Asian Journal of Research in Computer Science, 16, 128-148. https://doi.org/10.9734/AJRCOS/2023/v16i3351
Sathe, M. T., & Adamuthe, A. C. (2021). Comparative Study of Supervised Algorithms for Prediction of Students' Performance. International Journal of Modern Education and Computer Science, 13(1), 1-19. https://doi.org/10.5815/ijmecs.2021.01.01
Villegas-Ch, W., Govea, J., & Revelo-Tapia, S. (2023). Improving Student Retention in Institutions of Higher Education through Machine Learning: A Sustainable Approach. Sustainability, 15, 14512. https://doi.org/10.3390/su151914512
Yağcı, M. (2022). Educational data mining: prediction of students' academic performance using machine learning algorithms. Smart Learning Environments, 9, 11. https://doi.org/10.1186/s40561-022-00192-z
Copyright (c) 2026 Manda Rohandi, Mukhlisulfatih Latief, Mohamad Ilyas Abas (Author)

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.







