Przeskocz do nawigacji głównej Przeskocz do wyszukiwania Przeskocz do głównej treści

Comparative Analysis of Multivariate Mixture Models for Clustering Cancer Expression Data

Wyniki badań: Rozdział w książce/raport/materiał konferencyjnyWkład w konferencjęrecenzja

Abstrakt

Clustering gene expression data is essential for understanding tumor heterogeneity but is challenged by high dimensionality, noise, and class imbalance. We evaluated three Gaussian mixture model (GMM) based clustering methods: classical GMM, GMM with variance decomposition, and GMM with feature saliency on transcriptomic data from The Cancer Genome Atlas (TCGA) involving three cancer type pairs with varying molecular similarity and sample balance: breast invasive carcinoma vs uterine carcinosarcoma (BRCA vs UCS), lung adenocarcinoma vs lung squamous cell carcinoma (LUAD vs LUSC), and pancreatic adenocarcinoma vs sarcoma (PAAD vs SARC). Performance was assessed using Adjusted Rand Index, Fowlkes-Mallows Index, and Normalized Mutual Information.Results showed method-specific strengths influenced by dataset characteristics. Feature saliency clustering excelled in the highly imbalanced BRCA vs UCS pair by effectively downweighting irrelevant features. Variance decomposition performed best for the molecularly similar and balanced LUAD vs LUSC pair, capturing subtle expression differences. Classical GMM achieved the highest accuracy for the moderately imbalanced PAAD vs SARC pair. However, no single method consistently outperformed others across all datasets.This study highlights that while latent-variable mixture models are promising for transcriptomic clustering, their performance depends on data-specific factors such as imbalance and molecular similarity. Further work is needed to develop robust, scalable methods capable of adapting to diverse biological datasets.

Język oryginałuangielski
Tytuł publikacji goszczącejICBRA 2025 - Proceedings of the 12th International Conference on Bioinformatics Research and Applications
WydawcaAssociation for Computing Machinery, Inc
Strony88-92
Liczba stron5
ISBN (elektroniczny)9798400715808
Identyfikatory DOI
Status publikacjiOpublikowano - 22 gru 2025
Wydarzenie2025 12th International Conference on Bioinformatics Research and Applications, ICBRA 2025 - Prague, Republika Czeska
Czas trwania: 19 wrz 202521 wrz 2025

Seria publikacji

NazwaICBRA 2025 - Proceedings of the 12th International Conference on Bioinformatics Research and Applications

Konferencja

Konferencja2025 12th International Conference on Bioinformatics Research and Applications, ICBRA 2025
Kraj/TerytoriumRepublika Czeska
MiejscowośćPrague
Okres19/09/2521/09/25

Cele SDG ONZ

Ten wynik przyczynia się do realizacji następujących celów zrównoważonego rozwoju

  1. Cel 3 - Dobre zdrowie i dobre samopoczucie
    Cel 3 Dobre zdrowie i dobre samopoczucie

Obszary tematyczne ASJC Scopus

  • Biotechnologia
  • Genetyka
  • Sztuczna inteligencja
  • Zastosowania informatyki
  • Medycyna (różne)

Fingerprint

Zanurz się w tematy badawcze publikacji „Comparative Analysis of Multivariate Mixture Models for Clustering Cancer Expression Data”. Razem tworzą niepowtarzalny odcisk palca.

Cytuj to