Skip to main navigation Skip to search Skip to main content

Efficient system for clustering of dynamic document database

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

We describe in this paper, a system that groups, classifies and finds the latent semantic features in a database composed of a large number of documents. The database will be constantly growing as users who co-create it will be adding more and more new documents. Users require a system to provide them information, both about a specific document, and about the entire set of documents. This information includes statistical data about words in documents, information about aspects in which this words appears, classification, clustering, etc. To meet these expectations we propose using methods for searching for hidden patterns in multivariable data. We apply machine learning algorithms for data analysis, useful in identifying local patterns in multivariate data. We consider two different algorithms described in the literature (1) Probabilistic Latent Semantic Analysis Method [2] and (2) Nonnegative Matrix Factorization algorithm described in [4] and used in the text analysis system [1].

Original languageEnglish
Title of host publicationCooperative Design, Visualization, and Engineering - 8th International Conference, CDVE 2011, Proceedings
Pages186-189
Number of pages4
DOIs
Publication statusPublished - 2011
Event8th International Conference on Cooperative Design, Visualization, and Engineering, CDVE 2011 - Hong Kong, China
Duration: 11 Sept 201114 Sept 2011

Publication series

NameLecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
Volume6874 LNCS
ISSN (Print)0302-9743
ISSN (Electronic)1611-3349

Conference

Conference8th International Conference on Cooperative Design, Visualization, and Engineering, CDVE 2011
Country/TerritoryChina
CityHong Kong
Period11/09/1114/09/11

Keywords

  • NMF
  • classification
  • clustering
  • document database
  • semantic features

ASJC Scopus subject areas

  • Theoretical Computer Science
  • General Computer Science

Fingerprint

Dive into the research topics of 'Efficient system for clustering of dynamic document database'. Together they form a unique fingerprint.

Cite this