Przeskocz do nawigacji głównej Przeskocz do wyszukiwania Przeskocz do głównej treści

CLEANSE - Cluster-based Undersampling Method

  • Silesian University of Technology

Wyniki badań: Wkład do czasopismaArtykuł z konferencjirecenzja

6 Cytowania z bazy Scopus

Abstrakt

Class imbalance is a common problem with datasets relating to various areas of life. It causes many traditional machine learning algorithms to tend to misclassify minority samples as majority ones. Despite various studies, the class imbalance still remains a relevant problem for which no one-size-fits-all solution has been found. In this paper, an undersampling method based on clustering is presented. In the proposed approach K-means algorithm is used to cluster data. In homogenous "majority" clusters, i.e., clusters containing objects of only the majority class, objects within the specified distance from the center are removed. In the case of non-homogeneous clusters, objects located at the class decision boundary are removed using the KNN algorithm. As tests have shown, the clustering-based solution can improve classification quality. The results of experiments show that in many cases the proposed solution outperformed other undersampling techniques described in the literature.

Język oryginałuangielski
Strony (od–do)4541-4550
Liczba stron10
CzasopismoProcedia Computer Science
Tom225
Identyfikatory DOI
Status publikacjiOpublikowano - 2023
Wydarzenie27th International Conference on Knowledge Based and Intelligent Information and Engineering Sytems, KES 2023 - Athens, Grecja
Czas trwania: 6 wrz 20238 wrz 2023

Obszary tematyczne ASJC Scopus

  • Informatyka ogólna

Fingerprint

Zanurz się w tematy badawcze publikacji „CLEANSE - Cluster-based Undersampling Method”. Razem tworzą niepowtarzalny odcisk palca.

Cytowanie