Skip to main navigation Skip to search Skip to main content

Abstract

We have performed a comprehensive study dedicated to the analysis and comparative review of various unsupervised clustering algorithms for binary data analysis across a variety of datasets. We have compared the implementation of binary, model-based clustering algorithms corresponding to a multivariate Bernoulli Mixture Model with seven widely used reference algorithms based on distances between data vectors and aggregation/averaging operations. In our comparisons, we used both simulated and real datasets. The objective of the study was to draw conclusions that could inform the practical design of binary data analysis tools, particularly regarding the choice of algorithm, its initialization, and the influence of data binarization on the performance of the clustering procedure.

Original languageEnglish
Article number115341
JournalApplied Soft Computing
Volume200
DOIs
Publication statusPublished - Aug 2026

Keywords

  • Bernoulli trial
  • Binary data
  • Clustering
  • Expectation–maximization
  • Multivariate

ASJC Scopus subject areas

  • Software

Fingerprint

Dive into the research topics of 'Comparison of algorithms for unsupervised clustering of binary data'. Together they form a unique fingerprint.

Cite this