Skip to main navigation Skip to search Skip to main content

The role of feature selection in text mining in the process of discovering missing clinical annotations – Case study

  • Aleksander Płaczek
  • , Alicja Płuciennik
  • , Mirosław Pach
  • , Michał Jarząb
  • , Dariusz Mrozek
  • WASKO S.A.
  • Silesian University of Technology
  • Maria Sklodowska-Curie Institute of Oncology

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

1 Citation (Scopus)

Abstract

Vocabulary used by the doctors to describe the results of medical procedures changes alongside with the new standards. Text data, which is immediately understandable by the medical professional, is difficult to use in mass scale analysis. Extraction of data relevant to the given case, e.g. Bethesda class, means taking on the challenge of normalizing the freeform text and all the grammatical forms associated with it. This is particularly difficult in the Polish language where words change their form significantly according to their function in the sentence. We found common black-box methods for text mining inaccurate for this purpose. Here we described a word-frequency-based method for annotation of text data for Bethesda class extraction. We compared them with an algorithm based on a decision tree C4.5. We showed how important is the choice of the method and range of features to avoid conflicting classification. Proposed algorithms allowed to avoid the rule-base limitations.

Original languageEnglish
Title of host publicationBeyond Databases, Architectures and Structures. Paving the Road to Smart Data Processing and Analysis - 15th International Conference, BDAS 2019, Proceedings
EditorsStanisław Kozielski, Dariusz Mrozek, Paweł Kasprowski, Bożena Małysiak-Mrozek, Daniel Kostrzewa
PublisherSpringer Verlag
Pages248-262
Number of pages15
ISBN (Print)9783030190927
DOIs
Publication statusPublished - 2019
Event15th International Conference Beyond Databases, Architectures and Structures, BDAS 2019 - Ustroń, Poland
Duration: 28 May 201931 May 2019

Publication series

NameCommunications in Computer and Information Science
Volume1018
ISSN (Print)1865-0929
ISSN (Electronic)1865-0937

Conference

Conference15th International Conference Beyond Databases, Architectures and Structures, BDAS 2019
Country/TerritoryPoland
CityUstroń
Period28/05/1931/05/19

Keywords

  • Feature selection
  • Inverse document frequency
  • Text mining
  • Text tiding
  • Unstructured medical text

ASJC Scopus subject areas

  • General Computer Science
  • General Mathematics

Fingerprint

Dive into the research topics of 'The role of feature selection in text mining in the process of discovering missing clinical annotations – Case study'. Together they form a unique fingerprint.

Cite this