TY - GEN
T1 - Efficient system for clustering of dynamic document database
AU - Foszner, Pawel
AU - Gruca, Aleksandra
AU - Polanski, Andrzej
PY - 2011
Y1 - 2011
N2 - We describe in this paper, a system that groups, classifies and finds the latent semantic features in a database composed of a large number of documents. The database will be constantly growing as users who co-create it will be adding more and more new documents. Users require a system to provide them information, both about a specific document, and about the entire set of documents. This information includes statistical data about words in documents, information about aspects in which this words appears, classification, clustering, etc. To meet these expectations we propose using methods for searching for hidden patterns in multivariable data. We apply machine learning algorithms for data analysis, useful in identifying local patterns in multivariate data. We consider two different algorithms described in the literature (1) Probabilistic Latent Semantic Analysis Method [2] and (2) Nonnegative Matrix Factorization algorithm described in [4] and used in the text analysis system [1].
AB - We describe in this paper, a system that groups, classifies and finds the latent semantic features in a database composed of a large number of documents. The database will be constantly growing as users who co-create it will be adding more and more new documents. Users require a system to provide them information, both about a specific document, and about the entire set of documents. This information includes statistical data about words in documents, information about aspects in which this words appears, classification, clustering, etc. To meet these expectations we propose using methods for searching for hidden patterns in multivariable data. We apply machine learning algorithms for data analysis, useful in identifying local patterns in multivariate data. We consider two different algorithms described in the literature (1) Probabilistic Latent Semantic Analysis Method [2] and (2) Nonnegative Matrix Factorization algorithm described in [4] and used in the text analysis system [1].
KW - NMF
KW - classification
KW - clustering
KW - document database
KW - semantic features
UR - https://www.scopus.com/pages/publications/80052815131
U2 - 10.1007/978-3-642-23734-8_31
DO - 10.1007/978-3-642-23734-8_31
M3 - Conference contribution
AN - SCOPUS:80052815131
SN - 9783642237331
T3 - Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
SP - 186
EP - 189
BT - Cooperative Design, Visualization, and Engineering - 8th International Conference, CDVE 2011, Proceedings
T2 - 8th International Conference on Cooperative Design, Visualization, and Engineering, CDVE 2011
Y2 - 11 September 2011 through 14 September 2011
ER -