Welcome to Journal of University of Chinese Academy of Sciences,Today is

›› 2006, Vol. 23 ›› Issue (5): 640-646.DOI: 10.7523/j.issn.2095-6134.2006.5.012

• 论文 • Previous Articles     Next Articles

Quality Evaluation for Three Textual Document Clustering Algorithms

LIU Wu-Hua, LUO Tie-Jian, WANG Wen-Jie   

  1. Graduate University of Chinese Academy of Sciences, Beijing 100039
  • Received:1900-01-01 Revised:1900-01-01 Online:2006-09-15

Abstract: Textual document clustering is one of the effective approaches to establish a classification instance of huge textual document set. Clustering Validation or Quality Evaluation techniques can be used to assess the efficiency and effective of a clustering algorithm. This paper presents the quality evaluation criterions from outer and inner. Based on these criterions we take three typical textual document clustering algorithms for assessment with experiments. The comparison results show that STC(Suffix Tree Clustering) algorithm is better than k-Means and Ant-Based clustering algorithms. The better performance of STC algorithm comes from that it takes accounts the linguistic property when processing the documents. Ant-Based clustering algorithm’s performance variation is affected by the input variables. It is necessary to adopt linguistic properties to improve the Ant-Based text clustering’s performance.

Key words: Textual document clustering, Quality evaluate, Clustering validation, STC, Ant-Based Clustering, k-Means Clustering

CLC Number: