TY - GEN
T1 - On the Impact of Noisy Labels on Supervised Classification Models
AU - Dubel, Rafał
AU - Wijata, Agata M.
AU - Nalepa, Jakub
N1 - Publisher Copyright:
© 2023, The Author(s), under exclusive license to Springer Nature Switzerland AG.
PY - 2023
Y1 - 2023
N2 - The amount of data generated daily grows tremendously in virtually all domains of science and industry, and its efficient storage, processing and analysis pose significant practical challenges nowadays. To automate the process of extracting useful insights from raw data, numerous supervised machine learning algorithms have been researched so far. They benefit from annotated training sets which are fed to the training routine which elaborates a model that is further deployed for a specific task. The process of capturing real-world data may lead to acquring noisy observations, ultimately affecting the models trained from such data. The impact of the label noise is, however, under-researched, and the robustness of classic learners against such noise remains unclear. We tackle this research gap and not only thoroughly investigate the classification capabilities of an array of widely-adopted machine learning models over a variety of contamination scenarios, but also suggest new metrics that could be utilized to quantify such models’ robustness. Our extensive computational experiments shed more light on the impact of training set contamination on the operational behavior of supervised learners.
AB - The amount of data generated daily grows tremendously in virtually all domains of science and industry, and its efficient storage, processing and analysis pose significant practical challenges nowadays. To automate the process of extracting useful insights from raw data, numerous supervised machine learning algorithms have been researched so far. They benefit from annotated training sets which are fed to the training routine which elaborates a model that is further deployed for a specific task. The process of capturing real-world data may lead to acquring noisy observations, ultimately affecting the models trained from such data. The impact of the label noise is, however, under-researched, and the robustness of classic learners against such noise remains unclear. We tackle this research gap and not only thoroughly investigate the classification capabilities of an array of widely-adopted machine learning models over a variety of contamination scenarios, but also suggest new metrics that could be utilized to quantify such models’ robustness. Our extensive computational experiments shed more light on the impact of training set contamination on the operational behavior of supervised learners.
KW - Supervised machine learning
KW - binary classification
KW - label noise
KW - robustness
UR - https://www.scopus.com/pages/publications/85168660303
U2 - 10.1007/978-3-031-36021-3_8
DO - 10.1007/978-3-031-36021-3_8
M3 - Conference contribution
AN - SCOPUS:85168660303
SN - 9783031360206
T3 - Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
SP - 111
EP - 119
BT - Computational Science – ICCS 2023 - 23rd International Conference, Proceedings
A2 - Mikyška, Jiří
A2 - de Mulatier, Clélia
A2 - Krzhizhanovskaya, Valeria V.
A2 - Sloot, Peter M.A.
A2 - Paszynski, Maciej
A2 - Dongarra, Jack J.
PB - Springer Science and Business Media Deutschland GmbH
T2 - 23rd International Conference on Computational Science, ICCS 2023
Y2 - 3 July 2023 through 5 July 2023
ER -