Abstract
Class imbalance is a common problem with datasets relating to various areas of life. It causes many traditional machine learning algorithms to tend to misclassify minority samples as majority ones. Despite various studies, the class imbalance still remains a relevant problem for which no one-size-fits-all solution has been found. In this paper, an undersampling method based on clustering is presented. In the proposed approach K-means algorithm is used to cluster data. In homogenous "majority" clusters, i.e., clusters containing objects of only the majority class, objects within the specified distance from the center are removed. In the case of non-homogeneous clusters, objects located at the class decision boundary are removed using the KNN algorithm. As tests have shown, the clustering-based solution can improve classification quality. The results of experiments show that in many cases the proposed solution outperformed other undersampling techniques described in the literature.
| Original language | English |
|---|---|
| Pages (from-to) | 4541-4550 |
| Number of pages | 10 |
| Journal | Procedia Computer Science |
| Volume | 225 |
| DOIs | |
| Publication status | Published - 2023 |
| Event | 27th International Conference on Knowledge Based and Intelligent Information and Engineering Sytems, KES 2023 - Athens, Greece Duration: 6 Sept 2023 → 8 Sept 2023 |
Keywords
- K-means clustering
- classification
- imbalanced dataset
- undersampling
ASJC Scopus subject areas
- General Computer Science
Fingerprint
Dive into the research topics of 'CLEANSE - Cluster-based Undersampling Method'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver