TY - GEN
T1 - The role of feature selection in text mining in the process of discovering missing clinical annotations – Case study
AU - Płaczek, Aleksander
AU - Płuciennik, Alicja
AU - Pach, Mirosław
AU - Jarząb, Michał
AU - Mrozek, Dariusz
N1 - Publisher Copyright:
© Springer Nature Switzerland AG 2019.
PY - 2019
Y1 - 2019
N2 - Vocabulary used by the doctors to describe the results of medical procedures changes alongside with the new standards. Text data, which is immediately understandable by the medical professional, is difficult to use in mass scale analysis. Extraction of data relevant to the given case, e.g. Bethesda class, means taking on the challenge of normalizing the freeform text and all the grammatical forms associated with it. This is particularly difficult in the Polish language where words change their form significantly according to their function in the sentence. We found common black-box methods for text mining inaccurate for this purpose. Here we described a word-frequency-based method for annotation of text data for Bethesda class extraction. We compared them with an algorithm based on a decision tree C4.5. We showed how important is the choice of the method and range of features to avoid conflicting classification. Proposed algorithms allowed to avoid the rule-base limitations.
AB - Vocabulary used by the doctors to describe the results of medical procedures changes alongside with the new standards. Text data, which is immediately understandable by the medical professional, is difficult to use in mass scale analysis. Extraction of data relevant to the given case, e.g. Bethesda class, means taking on the challenge of normalizing the freeform text and all the grammatical forms associated with it. This is particularly difficult in the Polish language where words change their form significantly according to their function in the sentence. We found common black-box methods for text mining inaccurate for this purpose. Here we described a word-frequency-based method for annotation of text data for Bethesda class extraction. We compared them with an algorithm based on a decision tree C4.5. We showed how important is the choice of the method and range of features to avoid conflicting classification. Proposed algorithms allowed to avoid the rule-base limitations.
KW - Feature selection
KW - Inverse document frequency
KW - Text mining
KW - Text tiding
KW - Unstructured medical text
UR - https://www.scopus.com/pages/publications/85065903506
U2 - 10.1007/978-3-030-19093-4_19
DO - 10.1007/978-3-030-19093-4_19
M3 - Conference contribution
AN - SCOPUS:85065903506
SN - 9783030190927
T3 - Communications in Computer and Information Science
SP - 248
EP - 262
BT - Beyond Databases, Architectures and Structures. Paving the Road to Smart Data Processing and Analysis - 15th International Conference, BDAS 2019, Proceedings
A2 - Kozielski, Stanisław
A2 - Mrozek, Dariusz
A2 - Kasprowski, Paweł
A2 - Małysiak-Mrozek, Bożena
A2 - Kostrzewa, Daniel
PB - Springer Verlag
T2 - 15th International Conference Beyond Databases, Architectures and Structures, BDAS 2019
Y2 - 28 May 2019 through 31 May 2019
ER -