Examining Techniques to Solving Imbalanced Datasets in Educational Data Mining Systems

Research output: Contribution to journalArticlepeer-review

4 Citations (Scopus)

Abstract

The educational data mining research attempts have contributed in developing policies to improve student learning in different levels of educational institutions. One of the common challenges to building accurate classification and prediction systems is the imbalanced distribution of classes in the data collected. This study investigates data-level techniques and algorithm-level techniques. Six classifiers from each technique are used to explore their effectiveness to handle the imbalanced data problem while predicting students’ graduation grade based on their performance at the first stage. The classifiers are tested using the k-fold cross-validation approach before and after applying the data-level and algorithm-level techniques. For the purpose of evaluation, various evaluation metrics have been used such as accuracy, precision, recall, and f1-score. The results showed that the classifiers do not perform well with imbalanced dataset, and the performance could be improved by using these techniques. As for the level of improvement, it varies from one technique to another.

Original languageEnglish
Pages (from-to)205-213
Number of pages9
JournalInternational Journal of Computing
Volume21
Issue number2
DOIs
Publication statusPublished - Jun 30 2022

Keywords

  • Educational data mining
  • Imbalanced datasets
  • Machine learning
  • Prediction
  • Student grade

ASJC Scopus subject areas

  • Computer Science (miscellaneous)
  • Software
  • Information Systems
  • Hardware and Architecture
  • Computer Networks and Communications

Cite this