TY - GEN
T1 - Genetic selection of training sets for (not only) artificial neural networks
AU - Nalepa, Jakub
AU - Myller, Michal
AU - Piechaczek, Szymon
AU - Hrynczenko, Krzysztof
AU - Kawulok, Michal
N1 - Publisher Copyright:
© Springer Nature Switzerland AG 2018.
PY - 2018
Y1 - 2018
N2 - Creating high-quality training sets is the first step in designing robust classifiers. However, it is fairly difficult in practice when the data quality is questionable (data is heterogeneous, noisy and/or massively large). In this paper, we show how to apply a genetic algorithm for evolving training sets from data corpora, and exploit it for artificial neural networks (ANNs) alongside other state-of-the-art models. ANNs have been proved very successful in tackling a wide range of pattern recognition tasks. However, they suffer from several drawbacks, with selection of appropriate network topology and training sets being one of the most challenging in practice, especially when ANNs are trained using time-consuming back-propagation. Our experimental study (coupled with statistical tests), performed for both real-life and benchmark datasets, proved the applicability of a genetic algorithm to select training data for various classifiers which then generalize well to unseen data.
AB - Creating high-quality training sets is the first step in designing robust classifiers. However, it is fairly difficult in practice when the data quality is questionable (data is heterogeneous, noisy and/or massively large). In this paper, we show how to apply a genetic algorithm for evolving training sets from data corpora, and exploit it for artificial neural networks (ANNs) alongside other state-of-the-art models. ANNs have been proved very successful in tackling a wide range of pattern recognition tasks. However, they suffer from several drawbacks, with selection of appropriate network topology and training sets being one of the most challenging in practice, especially when ANNs are trained using time-consuming back-propagation. Our experimental study (coupled with statistical tests), performed for both real-life and benchmark datasets, proved the applicability of a genetic algorithm to select training data for various classifiers which then generalize well to unseen data.
KW - ANN
KW - Classification
KW - Genetic algorithm
KW - Training set selection
UR - https://www.scopus.com/pages/publications/85053845445
U2 - 10.1007/978-3-319-99987-6_15
DO - 10.1007/978-3-319-99987-6_15
M3 - Conference contribution
AN - SCOPUS:85053845445
SN - 9783319999869
T3 - Communications in Computer and Information Science
SP - 194
EP - 206
BT - Beyond Databases, Architectures and Structures. Facing the Challenges of Data Proliferation and Growing Variety - 14th International Conference, BDAS 2018, Held at the 24th IFIP World Computer Congress, WCC 2018, Proceedings
A2 - Kozielski, Stanislaw
A2 - Mrozek, Dariusz
A2 - Kasprowski, Pawel
A2 - Malysiak-Mrozek, Bozena
A2 - Kostrzewa, Daniel
PB - Springer Verlag
T2 - 14th International Conference on Beyond Databases, Architectures and Structures, BDAS 2018 Held at the 24th IFIP World Computer Congress, WCC 2018
Y2 - 18 September 2018 through 20 September 2018
ER -