Abstract
Discretisation is a processing step often included in the preliminary data preparation. Typically, when the input features have continuous domains and their discrete forms are needed, all are translated into categorical type at the same time, before data mining takes place. However, proceeding this way is not always the most advantageous to performance. The paper presents results from the research where the discretisation transformations were carried out sequentially forward for variables, and their selection was based on their values and also importance of the attributes estimated by the constructed rankings. The experiments were executed on the datasets from the area of stylometric analysis of texts, the application domain focused on recognising authorship based on individual characteristics of writing styles. For the selected data mining techniques, the performance was studied in the context of transformed features. The observed trends indicate that along with enhanced understanding of the nature of the data, partial discretisation of feature sets could bring higher accuracy than transformation of entire input domain, showing the merits of the described research methodology.
| Original language | English |
|---|---|
| Article number | 2679 |
| Journal | Applied Sciences (Switzerland) |
| Volume | 16 |
| Issue number | 6 |
| DOIs | |
| Publication status | Published - Mar 2026 |
Keywords
- classification
- discretisation
- ranking
- relevance
- stylometry
ASJC Scopus subject areas
- General Materials Science
- Instrumentation
- General Engineering
- Process Chemistry and Technology
- Computer Science Applications
- Fluid Flow and Transfer Processes
Fingerprint
Dive into the research topics of 'Does All or Nothing Always Work Best? In Search of Advantageous Representation of Attributes'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver