面对海量邮政寄递数据,现有的构建于关系数据库上的数据仓库系统在做数据分析时具有建设成本高、分析能力会遇到瓶颈等缺点。Hadoop具有高可扩展、高性能和低成本等优点,被广泛应用于大数据的存储和分析。基于对Hadoop开源框架的研究,设计邮政寄递大数据分析系统,并对该系统进行部分实现。结合邮政安监系统工程需求展开实验,得出大数据分析系统的性能参数,为后续工程建设提供依据。
Facing massive postal delivery data, the existing data warehouse system based on the traditional relational database has problems of high construction cost and analysis capacity bottleneck. Nowadays, Hadoop is widely used in large data storage and analysis, and it has the advantages of high scalability, high performance, and so on. On the basis of studies of the open source framework of Hadoop, combining with practical engineering project, we proposed a delivery data analysis system based on Hadoop. we implemented some parts of the system. We obtained the performance parameters of this system. The parameters can be widely used in future building of the project.
[1] 中国产业研究院. 2015-2022年中国电子商务市场全景调研及投资战略咨询报告[EB/OL].(2015)[2016-09-14]. http://www.chyxx.com/research/201510/349060.html.
[2] Pavlo A, Paulson E, Rasin A. A comparison of approaches to large-scale data analysis[C]//ACM. SIGMOD International Conference on Management of Data. Rhode Island: SIGMOD, 2009: 165-178.
[3] Gunther N, Puglia P, Tomasette K. Hadoop superlinear scalability[J]. Communications of the ACM, 2015, 58(4): 46-55.
[4] Apache Foundation. Apache Hadoop [EB/OL]. (2016-01-20) [2016-09-14]. https://wiki.apache.org/hadoop.
[5] Dean J, Ghemawat S. MapReduce: simplified data processing on large clusters[C]//Conference on Symposium on Opearting Systems Design & Implementation. USENIX Association, 2004: 107-113.
[6] Ghemawat S, Gobioff H, Leung S T. The Google file system[J]. Acm Sigops Operating Systems Review, 2003, 37(5): 29-43.
[7] Chang F, Dean J, Ghemawat S, et al. Bigtable: a distributed storage system for structured data[J]. ACM Transactions on Computer Systems, 2008, 26(2):205-218.
[8] Cohen J, Dolan B, Dunlap M, et al. MAD skills: new analysis practices for big data[J]. Proceedings of the Vldb Endowment, 2009, 2(2): 1 481-1 492.
[9] 覃雄派, 王会举, 杜小勇,等. 大数据分析:RDBMS与MapReduce的竞争与共生[J]. 计算机光盘软件与应用, 2013, 23(7): 55-56.
[10] Laurie B, Laurie P. Apache: the defini-tive guide[M]. 3rd ed. O'Reilly & Associates, 2005:14-60.
[11] 蔡斌,陈湘萍. Hadoop 技术内幕: 深入解析Hadoop Common和HDFS架构设计与实现原理[M]. 北京:机械工业出版社, 2013: 151-184.
[12] 董西成. Hadoop技术内幕: 深入解析MapReduce架构设计与实现原理: in-depth study of mapreduce[M]. 北京:机械工业出版社, 2013:228-240.
[13] 魏迪. 基于Hadoop的海量业务数据分析平台的设计与实现[D]. 北京:北京邮电大学, 2013.
[14] Thusoo A, Sarma J S, Jain N, et al. Hive: a warehousing solution over a map-reduce framework[J]. Proceedings of the Vldb Endowment, 2009, 2(2):1 626-1 629.
[15] Hoffman S. Apache flume: distributed log collection for Hadoop[M]. Birmingham: Packt Publishing, 2013: 24-88.
[16] Gupta S, Dutt N, Gupta R, et al. SPARK: a high-level synthesis framework for applying parallelizing compiler trans-formations[C]//International Conference on Vlsi Design. IEEE Computer Society, 2003:461-466.