Welcome to Journal of University of Chinese Academy of Sciences,Today is
Research Articles

An auto-configuration tool for heterogeneous Hadoop cluster

  • DAI Dong ,
  • ZHOU Xue-Hai ,
  • YANG Feng ,
  • WANG Chao
Expand
  • Suzhou Advanced Institution of USTC, Embedded System Lab, Suzhou 215123, Jiangsu, China; Computer Science and Technology College, USTC, Hefei 230027, China

Received date: 2010-07-19

  Revised date: 2010-09-06

  Online published: 2011-11-15

Abstract

The rapid development of cloud computing makes the heterogeneous cluster a hot research topic. One of the urgent problems in this field is how to configure the cloud computing software platform running on a heterogeneous cluster. We propose a new tool, which can configure each local server automatically according to its hardware parameters in a Hadoop cluster and configures each server individually by analyzing the heterogeneous features and the running history of the Hadoop cluster instead of the traditional uniform way. The simulation and experiment results show that this tool can obviously reduce the maintenance cost without any degradation of system performance and improve performance compared to the manual optimization. It is believed that this method is general and extensible and should be valuable for other cloud computing platforms.

Cite this article

DAI Dong , ZHOU Xue-Hai , YANG Feng , WANG Chao . An auto-configuration tool for heterogeneous Hadoop cluster[J]. Journal of University of Chinese Academy of Sciences, 2011 , 28(6) : 793 -800 . DOI: 10.7523/j.issn.2095-6134.2011.6.013

References


[1] Apache Software Foundation. The apache hadoop project.http://hadoop.apache.org/, as of 15/06/2009.

[2] Wikipedia. Heterogeneous_network.http://en.wikipedia.org/wiki/Heterogeneous_network.

[3] Bhandarkar M, Gogate S, Bhat V. Hadoop performance tuning: a case study. http://cloud.citris-uc.org/system/files/private/BerkeleyPerformanceTuning.pdf.

[4] Hadoop cluster setup.http://hadoop.apache.org/common/docs/current/cluster_setup.html.

[5] 闻新, 周露, 李东江, 等. MatLab模糊逻辑工具箱的分析与应用
[M]. 北京:科学出版社, 2001.

[6] Dean J, Ghemawat S. Mapreduce: simplified data processing on large clusters //Proc of OSDI. 2004: 137-150.

[7] Chang F, Dean J, Ghemawat S, et al. Bigtable: a distributed storage system for structured data
[J]. ACM Trans Comput Syst, 2008,26(2): 1-26.

[8] Boulon J, Konwinski A, Qi R, et al. Chukwa, a large-scale monitoring system //Cloud Computing and its Applications. Chicago, IL, October 2008:1-5.

[9] Thusoo A, Sarma J S, Jain N, et al. Hive-A warehousing solution over a map-reduce framework
[J]. PVLDB, 2009, 2(2):1626-1629.

[10] Hadoop. Powered by Hadoop. http://wiki.apache.org/hadoop/PoweredBy.

[11] Murthy A C. Speeding up Hadoop. http://developer.yahoo.com/blogs/ydn/posts/2009/09/hadoop_summit_speeding_up_hadoop/.

[12] Sharma S. Advanced Hadoop tuning optimization. http://www.slideshare.net/ImpetusInfo/ppt-on-advanced-hadoop-tuning-n-optimisation.

[13] Ghemawat S, Gobioff H, Leung S. The Google file system //Proceedings of the 19th ACM Symposium on Operating Systems Principles 2003. SOSP, 2003:29-43. DOI: 10.1145/945445.945450, URL:http://portal.acm.org/citation.cfm?id=945445.945450.

[14] Kapadia N H, Fortes J A B, Brodley C E. Predictive application-performance modeling in a computational grid environment //Proceedings of the IEEE International Symposium on High Performance Distributed Computing 1999. 1999. DOI: 10.1109/HPDC.1999.805281.

[15] Wang G, Butt A R, Pandey P, et al. A simulation approach to evaluating design decisions in MapReduce setups //Proceedings of the 17th Annual Meeting of the IEEE/ACM International Symposium on Modelling, Analysis and Simulation of Computer and Telecommunication Systems, MASCOTS 2009. DOI: 10.1109/MASCOT.2009.5366973.2009: 1-11.

[16] Malley O O, Murthy A C. Winning a 60 second dash with a yellow elephant . TR, Yahoo! Inc, 2009.

Outlines

/