TY - GEN
T1 - Improvement of Random Undersampling to Avoid Excessive Removal of Points from a Given Area of the Majority Class
AU - Bach, Małgorzata
AU - Werner, Aleksandra
N1 - Publisher Copyright:
© 2021, Springer Nature Switzerland AG.
PY - 2021
Y1 - 2021
N2 - In this paper we focus on class imbalance issue which often leads to sub-optimal performance of classifiers. Despite many attempts to solve this problem, there is still a need to look for better ones, which can overcome the limitations of known methods. For this reason we developed a new algorithm that in contrast to traditional random undersampling removes maximum k nearest neighbors of the samples which belong to the majority class. In such a way, there has been achieved not only the effect of reduction in size of the majority set but also the excessive removal of too many points from the given area has been successfully prevented. The conducted experiments are provided for eighteen imbalanced datasets, and confirm the usefulness of the proposed method to improve the results of the classification task, as compared to other undersampling methods. Non-parametric statistical tests show that these differences are usually statistically significant.
AB - In this paper we focus on class imbalance issue which often leads to sub-optimal performance of classifiers. Despite many attempts to solve this problem, there is still a need to look for better ones, which can overcome the limitations of known methods. For this reason we developed a new algorithm that in contrast to traditional random undersampling removes maximum k nearest neighbors of the samples which belong to the majority class. In such a way, there has been achieved not only the effect of reduction in size of the majority set but also the excessive removal of too many points from the given area has been successfully prevented. The conducted experiments are provided for eighteen imbalanced datasets, and confirm the usefulness of the proposed method to improve the results of the classification task, as compared to other undersampling methods. Non-parametric statistical tests show that these differences are usually statistically significant.
KW - Classification
KW - Imbalanced dataset
KW - K-Nearest Neighbors methods
KW - Sampling methods
KW - Undersampling
UR - https://www.scopus.com/pages/publications/85111388048
U2 - 10.1007/978-3-030-77967-2_15
DO - 10.1007/978-3-030-77967-2_15
M3 - Conference contribution
AN - SCOPUS:85111388048
SN - 9783030779665
T3 - Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
SP - 172
EP - 186
BT - Computational Science – ICCS 2021 - 21st International Conference, Proceedings
A2 - Paszynski, Maciej
A2 - Kranzlmüller, Dieter
A2 - Kranzlmüller, Dieter
A2 - Krzhizhanovskaya, Valeria V.
A2 - Dongarra, Jack J.
A2 - Sloot, Peter M.A.
A2 - Sloot, Peter M.A.
A2 - Sloot, Peter M.A.
PB - Springer Science and Business Media Deutschland GmbH
T2 - 21st International Conference on Computational Science, ICCS 2021
Y2 - 16 June 2021 through 18 June 2021
ER -