›› 2006, Vol. 23 ›› Issue (5): 640-646.DOI: 10.7523/j.issn.2095-6134.2006.5.012
• 论文 • Previous Articles Next Articles
LIU Wu-Hua, LUO Tie-Jian, WANG Wen-Jie
Received:
Revised:
Online:
Abstract: Textual document clustering is one of the effective approaches to establish a classification instance of huge textual document set. Clustering Validation or Quality Evaluation techniques can be used to assess the efficiency and effective of a clustering algorithm. This paper presents the quality evaluation criterions from outer and inner. Based on these criterions we take three typical textual document clustering algorithms for assessment with experiments. The comparison results show that STC(Suffix Tree Clustering) algorithm is better than k-Means and Ant-Based clustering algorithms. The better performance of STC algorithm comes from that it takes accounts the linguistic property when processing the documents. Ant-Based clustering algorithm’s performance variation is affected by the input variables. It is necessary to adopt linguistic properties to improve the Ant-Based text clustering’s performance.
Key words: Textual document clustering, Quality evaluate, Clustering validation, STC, Ant-Based Clustering, k-Means Clustering
CLC Number:
TP181
LIU Wu-Hua, LUO Tie-Jian, WANG Wen-Jie. Quality Evaluation for Three Textual Document Clustering Algorithms[J]. , 2006, 23(5): 640-646.
0 / / Recommend
Add to citation manager EndNote|Ris|BibTeX
URL: http://journal.ucas.ac.cn/EN/10.7523/j.issn.2095-6134.2006.5.012
http://journal.ucas.ac.cn/EN/Y2006/V23/I5/640