欢迎访问中国科学院大学学报,今天是
电子信息与计算机科学

一种基于双控节点的Ceph写性能优化方法

  • 黄遵祥 ,
  • 朱磊基 ,
  • 熊勇
展开
  • 中国科学院上海微系统与信息技术研究所 中国科学院无线传感网与通信重点实验室, 上海 201800;
    中国科学院大学, 北京 100049

收稿日期: 2020-07-13

  修回日期: 2020-11-11

  网络出版日期: 2021-06-03

基金资助

全军共用信息系统装备预研专用技术项目(31511030302)资助

A Ceph write performance optimization method based on double-control nodes

  • HUANG Zunxiang ,
  • ZHU Leiji ,
  • XIONG Yong
Expand
  • CAS Key Lab of Wireless Sensor Network and Communication, Shanghai Institute of Microsystem and Information Technology, Chinese Academy of Sciences, Shanghai 201800, China;
    University of Chinese Academy of Sciences, Beijing 100049, China

Received date: 2020-07-13

  Revised date: 2020-11-11

  Online published: 2021-06-03

摘要

分布式存储系统Ceph由于采用多副本强一致性写入机制,造成集群写性能不理想。针对该问题,提出一种基于双控节点的Ceph写性能优化方法,首先利用双控双存储阵列节点,当一个控制器出现故障时,该节点中的另一个伙伴控制器创建新的OSD进程并快速接管故障控制器的存储阵列,从而保证数据存储的安全性和高可靠性,同时将写入机制优化为主副本OSD在本地写入日志盘后,就向客户端返回写完成,之后写入数据盘和其余从副本的完成情况则由主副本OSD继续收集并完成后续各类回调操作,从而降低非必要写操作对集群写性能的影响。最后对数据可用性和集群写性能进行实验测试,其中写性能测试分别从写延迟、吞吐量和IOPS等3个角度,对优化后的方法和Ceph原生写入机制在顺序写和随机写两方面进行比较,进一步验证优化方法在维护数据高可用的同时,对写性能提升的效果。

本文引用格式

黄遵祥 , 朱磊基 , 熊勇 . 一种基于双控节点的Ceph写性能优化方法[J]. 中国科学院大学学报, 2022 , 39(6) : 817 -826 . DOI: 10.7523/j.ucas.2021.0051

Abstract

Because the distributed storage system Ceph uses a multi-copy strong consistency write mechanism, the cluster write performance is not ideal. To solve this problem, this paper proposes a Ceph write optimization method based on double-control nodes. With double-control double-RAID nodes, when one controller fails, another partner controller in the node creates a new OSD process and quickly takes over the RAID of the failed controller, thereby ensuring the safety and high reliability of data storage. At the same time, the write mechanism is optimized as follows: after the primary OSD is written to the journal, the write completion is returned to the client. After that, the primary OSD continues to collect the completion status of the write data disk and other slave copies, and then completes callback operations. Thereby reducing the impact of unnecessary write operations on the write performance of the cluster. Finally, the data availability and cluster write performance are tested experimentally. The write performance test compares the optimized method and Ceph’s native write mechanism in terms of sequential write and random write from three perspectives of write latency, throughput and IOPS. It further verifies the effect of the optimization method on improving write performance while maintaining high data availability.

参考文献

[1] Weil S. Ceph: reliable, scalable, and high-performance distributed storage[D]. California: University of California at Santa Cruz, 2007.
[2] Weil S A, Leung A W, Brandt S A,et al. RADOS: a scalable, reliable storage service for petabyte-scale storage clusters[C]// PDSW'07:Proceedings of the 2nd International Workshop on Petascale Data Storage held in conjunction with Supercomputing '07. November 11, 2007. Reno, Nevada. New York: ACM Press, 2007: 35-44. DOI:10.1145/1374596.1374606.
[3] Weil S A, Brandt S A, Miller E L, et al. Ceph: a scalable, high-performance distributed file system[C]//OSDI'06: Proceedings of the 7th USENIX Symposlum on Operating Systems Design and Implementation. November 2006, California: USENIX Association, 2006: 307-320.
[4] Weil S A, Brandt S A, Miller E L, et al. CRUSH: controlled, scalable, decentralized placement of replicated data[C]// SC'06: Proceedings of the 2006 ACM/IEEE Conference on Supercomputing. November 11-17, 2006, Tampa, FL, USA. IEEE, 2006: 1-12. DOI:10.1109/SC.2006.19.
[5] Weil S A, Pollack K T, Brandt S A, et al. Dynamic metadata management for petabyte-scale file systems[C]//SC'04: Proceedings of the 2004 ACM/IEEE Conference on Supercomputing. November 6-12, 2004, Pittsburgh, PA, USA. IEEE, 2004: 4. DOI:10.1109/SC.2004.22.
[6] Oh M, Eom J, Yoon J, et al. Performance optimization for all flash scale-out storage[C]// 2016 IEEE International Conference on Cluster Computing. September 12-16, 2016, Taipei, Taiwan, China. IEEE, 2016: 316-325. DOI:10.1109/CLUSTER.2016.11.
[7] Brim M J, Dillow D A, Oral S, et al. Asynchronous object storage with QoS for scientific and commercial big data[C]//PDSW'13: Proceedings of the 8th Parallel Data Storage Workshop. November 2013. Denver Colorado. New York, NY, USA: ACM, 2013: 7-13. DOI:10.1145/2538542.2538565.
[8] Zhang X, Wang Y Q, Wang Q, et al. A new approach to double I/O performance for ceph distributed file system in cloud computing[C]// 2019 2nd International Conference on Data Intelligence and Security (ICDIS). June 28-30, 2019, South Padre Island, TX, USA. IEEE, 2019: 68-75. DOI:10.1109/ICDIS.2019.00018.
[9] 刘鑫伟. 基于Ceph分布式存储系统副本一致性研究[D].武汉:华中科技大学,2016.
[10] 姚朋成. Ceph异构存储优化机制研究[D].重庆:重庆邮电大学,2019.
[11] Zhang J Y, Wu Y W, Chung Y C. PROAR: a weak consistency model for ceph[C]//2016 IEEE 22nd International Conference on Parallel and Distributed Systems. December 13-16, 2016, Wuhan, China. IEEE, 2016: 347-353. DOI:10.1109/ICPADS.2016.0054.
[12] Hilmi M, Mulyana E, Hendrawan H, et al. Analysis of network capacity effect on ceph based cloud storage performance[C]// 2019 IEEE 13th International Conference on Telecommunication Systems, Services, and Applications. October 3-4, 2019, Bali, Indonesia. IEEE, 2019: 22-24. DOI:10.1109/TSSA48701.2019.8985455.
[13] Bani Yusuf I N, Mulyana E, Hendrawan H, et al. Utilizing CRUSH algorithm on ceph to build a cluster of reliable data storage[C]// 2019 IEEE 13th International Conference on Telecommunication Systems, Services, and Applications. October 3-4, 2019, Bali, Indonesia. IEEE, 2019: 17-21. DOI:10.1109/TSSA48701.2019.8985481.
[14] Zhan K, Xu L L, Yuan Z M, et al. Performance optimization of large files writes to ceph based on multiple pipelines algorithm[C]// 2018 IEEE Intl Conf on Parallel & Distributed Processing with Applications, Ubiquitous Computing & Communications, Big Data & Cloud Computing, Social Computing & Networking, Sustainable Computing & Communications. December 11-13, 2018, Melbourne, VIC, Australia. IEEE, 2018: 525-532. DOI:10.1109/BDCloud.2018.00084.
[15] Fan Y, Wang Y, Ye M. An improved small file storage strategy in ceph file system[C]// 2018 14th International Conference on Computational Intelligence and Security (CIS). November 16-19, 2018, Hangzhou, China. IEEE, 2018: 488-491. DOI:10.1109/CIS2018.2018.00116.
[16] 邵曦煜, 李京, 周志强. 一种Ceph块设备跨集群迁移算法[J]. 中国科学技术大学学报, 2018, 48(9): 748-754. DOI:10.3969/j.issn.0253-2778.2018.09.009.
[17] Oh M, Park S, Yoon J, et al. Design of global data deduplication for a scale-out distributed storage system[C]// 2018 IEEE 38th International Conference on Distributed Computing Systems. July 2-6, 2018, Vienna, Austria. IEEE, 2018: 1063-1073. DOI:10.1109/ICDCS.2018.00106.
[18] Zhan K, Piao A H. Optimization of ceph reads/writes based on multi-threaded algorithms[C]// 2016 IEEE 18th International Conference on High Performance Computing and Communications; IEEE 14th International Conference on Smart City; IEEE 2nd International Conference on Data Science and Systems. December 12-14, 2016, Sydney, NSW, Australia. IEEE, 2016: 719-725. DOI:10.1109/HPCC-SmartCity-DSS.2016.0105.
文章导航

/