Skip to main navigation Skip to search Skip to main content

The class imbalance problem in construction of training datasets for authorship attribution

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

14 Citations (Scopus)

Abstract

The paper presents research on class imbalance in the context of construction of training sets for authorship recognition. In experiments the sets are artificially imbalanced, then balanced by under-sampling and over-sampling. The prepared sets are used in learning of two predictors: connectionist and rule-based, and their performance observed. The tests show that for artificial neural networks in several cases the predictive accuracy is not degraded but in fact improved, while one rule classifier is highly sensitive to class balance as it never performs better than for the original balanced set and in many cases worse.

Original languageEnglish
Title of host publicationMan–Machine Interactions - 4th International Conference on Man–Machine Interactions, ICMMI 2015
EditorsTadeusz Czachórski, Aleksandra Gruca, Agnieszka Brachman, Stanisław Kozielski, Tadeusz Czachórski
PublisherSpringer Verlag
Pages535-547
Number of pages13
ISBN (Print)9783319234366
DOIs
Publication statusPublished - 2016
Event4th International Conference on Man–Machine Interactions, ICMMI 2015 - Kocierz Pass, Poland
Duration: 6 Oct 20159 Oct 2015

Publication series

NameAdvances in Intelligent Systems and Computing
Volume391
ISSN (Print)2194-5357

Conference

Conference4th International Conference on Man–Machine Interactions, ICMMI 2015
Country/TerritoryPoland
CityKocierz Pass
Period6/10/159/10/15

Keywords

  • Authorship attribution
  • Class imbalance
  • Sampling strategy

ASJC Scopus subject areas

  • Control and Systems Engineering
  • General Computer Science

Fingerprint

Dive into the research topics of 'The class imbalance problem in construction of training datasets for authorship attribution'. Together they form a unique fingerprint.

Cite this