Because the distributed storage system Ceph uses a multi-copy strong consistency write mechanism, the cluster write performance is not ideal. To solve this problem, this paper proposes a Ceph write optimization method based on double-control nodes. With double-control double-RAID nodes, when one controller fails, another partner controller in the node creates a new OSD process and quickly takes over the RAID of the failed controller, thereby ensuring the safety and high reliability of data storage. At the same time, the write mechanism is optimized as follows: after the primary OSD is written to the journal, the write completion is returned to the client. After that, the primary OSD continues to collect the completion status of the write data disk and other slave copies, and then completes callback operations. Thereby reducing the impact of unnecessary write operations on the write performance of the cluster. Finally, the data availability and cluster write performance are tested experimentally. The write performance test compares the optimized method and Ceph’s native write mechanism in terms of sequential write and random write from three perspectives of write latency, throughput and IOPS. It further verifies the effect of the optimization method on improving write performance while maintaining high data availability.
HUANG Zunxiang
,
ZHU Leiji
,
XIONG Yong
. A Ceph write performance optimization method based on double-control nodes[J]. Journal of University of Chinese Academy of Sciences, 2022
, 39(6)
: 817
-826
.
DOI: 10.7523/j.ucas.2021.0051
[1] Weil S. Ceph: reliable, scalable, and high-performance distributed storage[D]. California: University of California at Santa Cruz, 2007.
[2] Weil S A, Leung A W, Brandt S A,et al. RADOS: a scalable, reliable storage service for petabyte-scale storage clusters[C]// PDSW'07:Proceedings of the 2nd International Workshop on Petascale Data Storage held in conjunction with Supercomputing '07. November 11, 2007. Reno, Nevada. New York: ACM Press, 2007: 35-44. DOI:10.1145/1374596.1374606.
[3] Weil S A, Brandt S A, Miller E L, et al. Ceph: a scalable, high-performance distributed file system[C]//OSDI'06: Proceedings of the 7th USENIX Symposlum on Operating Systems Design and Implementation. November 2006, California: USENIX Association, 2006: 307-320.
[4] Weil S A, Brandt S A, Miller E L, et al. CRUSH: controlled, scalable, decentralized placement of replicated data[C]// SC'06: Proceedings of the 2006 ACM/IEEE Conference on Supercomputing. November 11-17, 2006, Tampa, FL, USA. IEEE, 2006: 1-12. DOI:10.1109/SC.2006.19.
[5] Weil S A, Pollack K T, Brandt S A, et al. Dynamic metadata management for petabyte-scale file systems[C]//SC'04: Proceedings of the 2004 ACM/IEEE Conference on Supercomputing. November 6-12, 2004, Pittsburgh, PA, USA. IEEE, 2004: 4. DOI:10.1109/SC.2004.22.
[6] Oh M, Eom J, Yoon J, et al. Performance optimization for all flash scale-out storage[C]// 2016 IEEE International Conference on Cluster Computing. September 12-16, 2016, Taipei, Taiwan, China. IEEE, 2016: 316-325. DOI:10.1109/CLUSTER.2016.11.
[7] Brim M J, Dillow D A, Oral S, et al. Asynchronous object storage with QoS for scientific and commercial big data[C]//PDSW'13: Proceedings of the 8th Parallel Data Storage Workshop. November 2013. Denver Colorado. New York, NY, USA: ACM, 2013: 7-13. DOI:10.1145/2538542.2538565.
[8] Zhang X, Wang Y Q, Wang Q, et al. A new approach to double I/O performance for ceph distributed file system in cloud computing[C]// 2019 2nd International Conference on Data Intelligence and Security (ICDIS). June 28-30, 2019, South Padre Island, TX, USA. IEEE, 2019: 68-75. DOI:10.1109/ICDIS.2019.00018.
[9] 刘鑫伟. 基于Ceph分布式存储系统副本一致性研究[D].武汉:华中科技大学,2016.
[10] 姚朋成. Ceph异构存储优化机制研究[D].重庆:重庆邮电大学,2019.
[11] Zhang J Y, Wu Y W, Chung Y C. PROAR: a weak consistency model for ceph[C]//2016 IEEE 22nd International Conference on Parallel and Distributed Systems. December 13-16, 2016, Wuhan, China. IEEE, 2016: 347-353. DOI:10.1109/ICPADS.2016.0054.
[12] Hilmi M, Mulyana E, Hendrawan H, et al. Analysis of network capacity effect on ceph based cloud storage performance[C]// 2019 IEEE 13th International Conference on Telecommunication Systems, Services, and Applications. October 3-4, 2019, Bali, Indonesia. IEEE, 2019: 22-24. DOI:10.1109/TSSA48701.2019.8985455.
[13] Bani Yusuf I N, Mulyana E, Hendrawan H, et al. Utilizing CRUSH algorithm on ceph to build a cluster of reliable data storage[C]// 2019 IEEE 13th International Conference on Telecommunication Systems, Services, and Applications. October 3-4, 2019, Bali, Indonesia. IEEE, 2019: 17-21. DOI:10.1109/TSSA48701.2019.8985481.
[14] Zhan K, Xu L L, Yuan Z M, et al. Performance optimization of large files writes to ceph based on multiple pipelines algorithm[C]// 2018 IEEE Intl Conf on Parallel & Distributed Processing with Applications, Ubiquitous Computing & Communications, Big Data & Cloud Computing, Social Computing & Networking, Sustainable Computing & Communications. December 11-13, 2018, Melbourne, VIC, Australia. IEEE, 2018: 525-532. DOI:10.1109/BDCloud.2018.00084.
[15] Fan Y, Wang Y, Ye M. An improved small file storage strategy in ceph file system[C]// 2018 14th International Conference on Computational Intelligence and Security (CIS). November 16-19, 2018, Hangzhou, China. IEEE, 2018: 488-491. DOI:10.1109/CIS2018.2018.00116.
[16] 邵曦煜, 李京, 周志强. 一种Ceph块设备跨集群迁移算法[J]. 中国科学技术大学学报, 2018, 48(9): 748-754. DOI:10.3969/j.issn.0253-2778.2018.09.009.
[17] Oh M, Park S, Yoon J, et al. Design of global data deduplication for a scale-out distributed storage system[C]// 2018 IEEE 38th International Conference on Distributed Computing Systems. July 2-6, 2018, Vienna, Austria. IEEE, 2018: 1063-1073. DOI:10.1109/ICDCS.2018.00106.
[18] Zhan K, Piao A H. Optimization of ceph reads/writes based on multi-threaded algorithms[C]// 2016 IEEE 18th International Conference on High Performance Computing and Communications; IEEE 14th International Conference on Smart City; IEEE 2nd International Conference on Data Science and Systems. December 12-14, 2016, Sydney, NSW, Australia. IEEE, 2016: 719-725. DOI:10.1109/HPCC-SmartCity-DSS.2016.0105.