Skip to main navigation Skip to search Skip to main content

Comparative Analysis of Multivariate Mixture Models for Clustering Cancer Expression Data

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Clustering gene expression data is essential for understanding tumor heterogeneity but is challenged by high dimensionality, noise, and class imbalance. We evaluated three Gaussian mixture model (GMM) based clustering methods: classical GMM, GMM with variance decomposition, and GMM with feature saliency on transcriptomic data from The Cancer Genome Atlas (TCGA) involving three cancer type pairs with varying molecular similarity and sample balance: breast invasive carcinoma vs uterine carcinosarcoma (BRCA vs UCS), lung adenocarcinoma vs lung squamous cell carcinoma (LUAD vs LUSC), and pancreatic adenocarcinoma vs sarcoma (PAAD vs SARC). Performance was assessed using Adjusted Rand Index, Fowlkes-Mallows Index, and Normalized Mutual Information.Results showed method-specific strengths influenced by dataset characteristics. Feature saliency clustering excelled in the highly imbalanced BRCA vs UCS pair by effectively downweighting irrelevant features. Variance decomposition performed best for the molecularly similar and balanced LUAD vs LUSC pair, capturing subtle expression differences. Classical GMM achieved the highest accuracy for the moderately imbalanced PAAD vs SARC pair. However, no single method consistently outperformed others across all datasets.This study highlights that while latent-variable mixture models are promising for transcriptomic clustering, their performance depends on data-specific factors such as imbalance and molecular similarity. Further work is needed to develop robust, scalable methods capable of adapting to diverse biological datasets.

Original languageEnglish
Title of host publicationICBRA 2025 - Proceedings of the 12th International Conference on Bioinformatics Research and Applications
PublisherAssociation for Computing Machinery, Inc
Pages88-92
Number of pages5
ISBN (Electronic)9798400715808
DOIs
Publication statusPublished - 22 Dec 2025
Event2025 12th International Conference on Bioinformatics Research and Applications, ICBRA 2025 - Prague, Czech Republic
Duration: 19 Sept 202521 Sept 2025

Publication series

NameICBRA 2025 - Proceedings of the 12th International Conference on Bioinformatics Research and Applications

Conference

Conference2025 12th International Conference on Bioinformatics Research and Applications, ICBRA 2025
Country/TerritoryCzech Republic
CityPrague
Period19/09/2521/09/25

UN SDGs

This output contributes to the following UN Sustainable Development Goals (SDGs)

  1. SDG 3 - Good Health and Well-being
    SDG 3 Good Health and Well-being

Keywords

  • Expectation-Maximization
  • feature saliency
  • Gaussian Mixture Model
  • unsupervised learning
  • variance decomposition

ASJC Scopus subject areas

  • Biotechnology
  • Genetics
  • Artificial Intelligence
  • Computer Science Applications
  • Medicine (miscellaneous)

Fingerprint

Dive into the research topics of 'Comparative Analysis of Multivariate Mixture Models for Clustering Cancer Expression Data'. Together they form a unique fingerprint.

Cite this