Journal of University of Chinese Academy of Sciences >
An XBRL dimensional data parsing algorithm based on the Map/Reduce parallel programming model
Received date: 2013-04-26
Revised date: 2013-05-20
Online published: 2014-01-15
This article intends to study mass semi-structured data processing technology from XBRL dimensional data processing perspective. A new XBRL dimensional data parsing algorithm is proposed based on the Map/Reduce parallel programming model and StAX stream parsing technique. The algorithm specifically targets the analysis of complex data reference relationships among XML files in the XBRL financial report. In order to parse complex XBRL dimensional data, the algorithm uses a single XBRL financial report as the minimum processing unit. First, the data are extracted from the dimensional fact items, and then the business semantic data are processed. In experimental tests, the proposed algorithm presents an obvious advantage in large-scale XBRL data processing.
ZHU Jianpeng , WANG Ying , YANG Cheng . An XBRL dimensional data parsing algorithm based on the Map/Reduce parallel programming model[J]. Journal of University of Chinese Academy of Sciences, 2014 , 31(1) : 124 -129 . DOI: 10.7523/j.issn.2095-6134.2014.01.018
[1] Li G J, Cheng X Q. Research status and scientific thinking of big data[J]. Bulletin of Chinese Academy of Sciences, 2012, 27(6):647-657(in Chinese). 李国杰, 程学旗. 大数据研究:未来科技及经济社会发展的重大战略领域:大数据的研究现状与科学思考[J]. 中国科学院院刊, 2012, 27(6):647-657.
[2] Shi Z Z. Big data mining in the cloud, intelligent information processing VI[M]. Springer Berlin Heidelberg, 2012: 13-14.
[3] Mika S I. Preface to part Ⅲ adaptive big data analytics. procedia computer science[M]. Elsevier B V, 2013: 211.
[4] Jeffrey D. MapReduce: a flexible data processing tool[J]. Communications of the ACM, 2010, 53(1): 72-77.
[5] Qin X P, Wang H J, Du X Y, et al. Big data analysis: competition and symbiosis of RDBMS and MapReduce[J]. Journal of Software, 2012, 23(1): 32-45(in Chinese). 覃雄派, 王会举, 杜小勇, 等. 大数据分析:RDBMS与MapReduce的竞争与共生[J]. 软件学报, 2012, 23(1): 32-45.
[6] Dean J, Ghemawat S. MapReduce: simplified data processing on large clusters[J]. Communications of the ACM, 2008, 51(1): 107-113.
[7] Michele T, Stefano Crespi-Reghizzi. Parallel iterative compilation: using MapReduce to speedup machine learning in compilers[C]//The Third International Workshop on MapReduce and its Applications (MAPREDUCE'12). ACM New York, NY, USA, 2012:18-19.
[8] Daniel Z, Shawn B, Sven K, et al. Parallelizing XML data-streaming workflows via MapReduce[J]. Journal of Computer and System Sciences, 2010, 76(6):447-463.
/
| 〈 |
|
〉 |