Abstrakt
Class imbalance is a common problem with datasets relating to various areas of life. It causes many traditional machine learning algorithms to tend to misclassify minority samples as majority ones. Despite various studies, the class imbalance still remains a relevant problem for which no one-size-fits-all solution has been found. In this paper, an undersampling method based on clustering is presented. In the proposed approach K-means algorithm is used to cluster data. In homogenous "majority" clusters, i.e., clusters containing objects of only the majority class, objects within the specified distance from the center are removed. In the case of non-homogeneous clusters, objects located at the class decision boundary are removed using the KNN algorithm. As tests have shown, the clustering-based solution can improve classification quality. The results of experiments show that in many cases the proposed solution outperformed other undersampling techniques described in the literature.
| Język oryginału | angielski |
|---|---|
| Strony (od–do) | 4541-4550 |
| Liczba stron | 10 |
| Czasopismo | Procedia Computer Science |
| Tom | 225 |
| Identyfikatory DOI | |
| Status publikacji | Opublikowano - 2023 |
| Wydarzenie | 27th International Conference on Knowledge Based and Intelligent Information and Engineering Sytems, KES 2023 - Athens, Grecja Czas trwania: 6 wrz 2023 → 8 wrz 2023 |
Obszary tematyczne ASJC Scopus
- Informatyka ogólna
Fingerprint
Zanurz się w tematy badawcze publikacji „CLEANSE - Cluster-based Undersampling Method”. Razem tworzą niepowtarzalny odcisk palca.Cytowanie
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver