欢迎访问中国科学院大学学报,今天是
论文

一种基于密度最大值的聚类算法

  • 王晶 ,
  • 夏鲁宁 ,
  • 荆继武
展开
  • 1. 中国科学技术大学电子工程与信息科学系, 合肥 230027;
    2. 中国科学院研究生院信息安全国家重点实验室, 北京 100049

收稿日期: 2008-10-08

  修回日期: 2009-01-09

  网络出版日期: 2009-07-15

基金资助

国家863计划(2006AA01Z454)和电子信息产业发展基金资助 

Maximum density clustering algorithm

  • WANG Jing ,
  • XIA Lu-Ning ,
  • JING Ji-Wu
Expand
  • 1. Department of Electronic Engineering and Information Science, University of Science and Technology of China, Hefei 230027, China;
    2. State Key Lab of Information Security, Graduate University of the Chinese Academy of Sciences, Beijing 100049, China

Received date: 2008-10-08

  Revised date: 2009-01-09

  Online published: 2009-07-15

摘要

提出了一种结合了基于密度聚类思想的划分聚类方法——"密度最大值聚类算法(MDCA)",以最大密度对象作为起始点,通过考察最大密度对象所处空间区域的密度分布情况来划分基本簇,并合并基本簇获得最终的簇划分.实验表明,MDCA能够自动确定簇数量,并有效发现任意形状的簇,对于未知数据集的处理能力和聚类准确度都优于传统的基于划分聚类算法.

本文引用格式

王晶 , 夏鲁宁 , 荆继武 . 一种基于密度最大值的聚类算法[J]. 中国科学院大学学报, 2009 , 26(4) : 539 -548 . DOI: 10.7523/j.issn.2095-6134.2009.4.016

Abstract

This paper proposes a new clustering algorithm named maximum density clustering algorithm(MDCA). In MDCA the concept of density is introduced to identify the count of clusters automatically.By selecting the densest object as the threshold, densities of those objects around the densest object are reviewed to decide the partition of basic blocks. Then the basic blocks are merged to form clusters of arbitrary shape. Experiments show that the ability and validity of MDCA in processing unknown datasets are all better than traditional partition-based clustering algorithms.

参考文献


[1] MacQueen J. Some methods for classification and analysis of multivariate observations //LeCam L M,Neyman J,eds. Proc of Fifth Berkeley Symposium on Math. Stat and Prob: University of California Press, 1967:281-297.

[2] Tan P N,Steinbach M,等著. 范 明,范宏建,等译.数据挖掘导论(Introduction to Data Mining)
[M]. 北京:人民邮电出版社, 2006.

[3] Ester M, Kriegel H P, Sander J. A density-based algorithm for discovering clusters in large spatial databases with noise //Usama M Fayyad, Padhraic Smyth, Gregory Piatetsky-Shapiro,eds. Proc of 2nd International Conference on Knowledge Discovery and Data Mining (KDD'96). Portland: ACM Press, 1996:226-231.

[4] Ankerst M, Breunig M M, et al. OPTICS: ordering points to identify the clustering structure //Alex Delis, Christos Faloutsos, Shahram Ghandeharizadeh, eds. Proc ACM SIGMOD'99 Int Conf on Management of Data. Philadelphia Pennsylvania: ACM Press, 1999:49-60.

[5] Agrawal R, Gehrke J, Gunopulos D, et al. Automatic subspace clustering of high dimensional data for data mining applications //Laura Haas,Ashutosh Tiwary,eds. Proc of 1998 ACM-SIGMOD Intl Conf on Management of Data. Seattle, Washington: ACM Press, 1998:94-105.

[6] Katsavounidis I, Kuo C, Zhang Z. A new initialization technique for generalized lloyd iteration
[J]. IEEE Signal Processing Letters, 1994, 1(10): 144-146.

[7] Tou J T,Gonzalez R C. Pattern recognition principles
[M].Dyersburg, TN, USA: Addison-Wesley, 1975.

[8] Christian Mauceri, Diem Ho. Clustering by kernel density
[J]. Computational Economics, 2007, 29(2): 199-212.

[9] Liu N, Zhang B Y, Yan J, et al. Learning similarity measures in the non-orthogonal space //Grossman D, Gravano L, Zhai C, Herzog O, Evans D, eds. Proc of the 13th Conf on Information and Knowledge Management (CIKM 2004). New York: ACM Press, 2004:334-341.

[10] Jarvis R A, Patrick E A. Clustering using a similarity measure based on shared nearest neighbors
[J]. IEEE Transactions on Computers, 1973, C-22(11): 1025-1034.

[11] 王世儒.计算方法
[M]. 西安: 西安电子科技大学出版社,1999.

[12] Steinbach M, Karypis G, et al. A comparison of document clustering techniques. Computer Science and Engineering Technical Report, Report No. 00-034 . Minnesota USA: University of Minnesota, 2000.

文章导航

/