Abstract
Microarrays provide a new technique of measuring gene expression that attracted a lot of research interest in recent years. It has been suggested that gene expression data from microarrays (biochips) can be utilized in many biomedical areas, for example in cancer classification. Whereas several, new and existing, methods of classification has been tested, a selection of proper (optimal) set of genes, which expression serves during classification, is still an open problem. In this paper we propose a heuristic method of choosing suboptimal set of genes by using support vector machines (SVM). Obtained set of genes optimizes leave-one-out cross-validation error. The method is tested on microarray gene expression data of samples of two cancer types: acute myeloid leukemia (AML) and acute lymphoblastic leukemia (ALL). The results show that quality of classification is much better than for sets obtained using other methods of feature selection. In addition, we demonstrate that maximum separation in a training data set may lead to deterioration of performance in an independent validation data set, a phenomenon akin to overfitting.
| Original language | English |
|---|---|
| Pages (from-to) | 43-56 |
| Number of pages | 14 |
| Journal | Journal of Biological Systems |
| Volume | 11 |
| Issue number | 1 |
| DOIs | |
| Publication status | Published - 1 Mar 2003 |
UN SDGs
This output contributes to the following UN Sustainable Development Goals (SDGs)
-
SDG 3 Good Health and Well-being
Keywords
- Cancer diagnosis
- Classification
- Feature selection
- Gene expression data
- Support vector machines
ASJC Scopus subject areas
- Ecology
- Agricultural and Biological Sciences (miscellaneous)
- Applied Mathematics
Fingerprint
Dive into the research topics of 'A note on classification of gene expression data using support vector machines'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver