欢迎访问中国科学院大学学报,今天是
电子信息与计算机科学

一种改进YOLOv3的学校场所目标识别方法

  • 高锦风 ,
  • 陈玉 ,
  • 魏永明 ,
  • 李剑南 ,
  • 江若楠
展开
  • 1 中国科学院空天信息创新研究院, 北京 100094;
    2 中国科学院大学, 北京 100049;
    3 天津市城市规划设计研究总院有限公司 天津市智慧城市规划企业重点实验室, 天津 300000

收稿日期: 2021-12-01

  修回日期: 2021-12-20

  网络出版日期: 2021-12-20

基金资助

兵团科技攻关项目(2017DB005-01)资助

School place identification based on improved YOLOv3

  • GAO Jinfeng ,
  • CHEN Yu ,
  • WEI Yongming ,
  • LI Jiannan ,
  • JIANG Ruonan
Expand
  • 1 Aerospace Information Research Institute, Chinese Academy of Sciences, Beijing 100094, China;
    2 University of Chinese Academy of Sciences, Beijing 100049, China;
    3 Key Enterprises Laboratory of Smart City Planning of Tianjin, Tianjin Urban Planning and Design Institute Co., Ltd, Tianjin 300000, China

Received date: 2021-12-01

  Revised date: 2021-12-20

  Online published: 2021-12-20

摘要

基于遥感影像进行特定场所类型的识别在智慧城市规划、土地利用分析、平安城市建设等多方面都具有重要意义。然而,不同场所的环境景观属性(如道路和停车场等)比较复杂,难以用传统的分类或目标识别方法基于简单的规则进行识别。卷积神经网络具有较强的空间信息挖掘能力,尝试对著名的YOLOv3模型进行改进,提出一种名为YOLO-S-CIoU的新模型,用于学校场所目标的识别。主要改进工作包括:1)使用SRXnet模块替换YOLOv3中的Darknet53模块以提高特征学习能力;2)利用complete-IoU loss (CIoU loss)优化边界框的回归;3)基于自制的学校场所样本数据集(SS数据集)进行训练和验证。实验结果表明,YOLO-S-CIoU的平均精度(AP)达到96.46%;参数量为226 MB。与改进前YOLOv3相比,YOLO-S-CIoU实现了参数量9 MB的下降以及AP 2.3%的提升。此外,在新疆图木舒克市和烟台市区域遥感影像中对学校场所目标识别,召回率比YOLOv3分别提高37.5%和42.2%。这表明改进后的网络模型在不同地理区域的遥感影像识别中具有更强的鲁棒性和更高的识别能力。

本文引用格式

高锦风 , 陈玉 , 魏永明 , 李剑南 , 江若楠 . 一种改进YOLOv3的学校场所目标识别方法[J]. 中国科学院大学学报, 2023 , 40(4) : 531 -539 . DOI: 10.7523/j.ucas.2021.0081

Abstract

The specific place identification based on remote sensing images is of great significance in smart city planning, land use analysis, and safe city construction. However, it is difficult to identify these places with simple rules by using traditional classification or target recognition methods due to the complex environmental background. Convolutional neural networks have strong spatial information mining capabilities. In this article, we improved the well-known YOLOv3 model to a new model called YOLO-S-CIoU for school place identification. The main improvements include:1) Using the SRXnet module to replace the Darknet53 module in YOLOv3 to improve the feature learning ability; 2) Using complete-IoU loss (CIoU loss) to optimize the regression of the bounding box; 3) Training and verification based on self-made school place sample dataset (SS dataset). The results showed that the average accuracy (AP) of YOLO-S-CIoU reached 96.46%; the parameter amount was 226 MB. Compared with YOLOv3, YOLO-S-CIoU has achieved a 9 MB reduction in parameter volume and a 2.3% increase in AP. In addition, in the regional remote sensing images of Tumshuk and Yantai, the recall rates of school places were increased by 37.5% and 42.2% than YOLOv3, respectively. These works show that the improved model has stronger robustness and higher identification ability in remote sensing image identification in different areas.

参考文献

[1] 甄峰, 席广亮, 秦萧. 基于地理视角的智慧城市规划与建设的理论思考[J]. 地理科学进展, 2015, 34(4):402-409. DOI:10.11820/dlkxjz.2015.04.001.
[2] 黄海涛, 柯长青. 高分辨率遥感影像中操场跑道的自动提取[J]. 遥感信息, 2009, 24(3):19-22, 29. DOI:10.3969/j.issn.1000-3177.2009.03.005.
[3] 普恒. 高分遥感城市典型地物对象化识别方法研究[D]. 北京:北京建筑大学, 2019.
[4] 范荣双, 陈洋, 徐启恒, 等. 基于深度学习的高分辨率遥感影像建筑物提取方法[J]. 测绘学报, 2019, 48(1):34-41. DOI:CNKI:SUN:CHXB.0.2019-01-006.
[5] Wen Q, Jiang K Y, Wang W, et al. Automatic building extraction from google earth images under complex backgrounds based on deep instance segmentation network[J]. Sensors (Basel, Switzerland), 2019, 19(2):333. DOI:10.3390/s19020333.
[6] Chen Y, Wei Y M, Wang Q J, et al. Mapping post-earthquake landslide susceptibility:a U-net like approach[J]. Remote Sensing, 2020, 12(17):2767. DOI:10.3390/rs12172767.
[7] Girshick R, Donahue J, Darrell T, et al. Rich feature hierarchies for accurate object detection and semantic segmentation[C]//2014 IEEE Conference on Computer Vision and Pattern Recognition. June 23-28, 2014, Columbus, OH, USA. IEEE, 2014:580-587. DOI:10.1109/CVPR.2014.81.
[8] Girshick R. Fast R-CNN[C]//2015 IEEE International Conference on Computer Vision (ICCV). December 7-13, 2015, Santiago, Chile. IEEE, 2015:1440-1448. DOI:10.1109/ICCV.2015.169.
[9] Ren S Q, He K M, Girshick R, et al. Faster R-CNN:towards real-time object detection with region proposal networks[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2017, 39(6):1137-1149. DOI:10.1109/TPAMI.2016.2577031.
[10] He K M, Gkioxari G, Dollár P, et al. Mask R-CNN[C]//2017 IEEE International Conference on Computer Vision (ICCV). October 22-29, 2017, Venice, Italy. IEEE, 2017:2980-2988. DOI:10.1109/ICCV.2017.322.
[11] Redmon J, Divvala S, Girshick R, et al. You only look once:unified, real-time object detection[C]//2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). June 27-30, 2016, Las Vegas, NV, USA. IEEE, 2016:779-788. DOI:10.1109/CVPR.2016.91.
[12] Redmon J, Farhadi A. YOLO9000:better, faster, stronger[C]//2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). July 21-26, 2017, Honolulu, HI, USA. IEEE, 2017:6517-6525. DOI:10.1109/CVPR.2017.690.
[13] Redmon J, Farhadi A. YOLOv3:an incremental improvement[EB/OL]. arXiv:1804.02767. (2018-04-08)[2021-12-15]. https://arxiv.org/abs/1804. 02767.
[14] Alganci U, Soydas M, Sertel E. Comparative research on deep learning approaches for airplane detection from very high-resolution satellite images[J]. Remote Sensing, 2020, 12(3):458. DOI:10.3390/rs12030458.
[15] 余东行, 张宁, 张保明, 等. 结合卷积神经网络与显著性特征的机场检测[J]. 测绘通报, 2019(7):44-49. DOI:10.13474/j.cnki.11-2246.2019.0216.
[16] Ma H J, Liu Y L, Ren Y H, et al. Detection of collapsed buildings in post-earthquake remote sensing images based on the improved YOLOv3[J]. Remote Sensing, 2019, 12(1):44. DOI:10.3390/rs12010044.
[17] 陈连凯, 李邦昱, 齐亮. 融合图像显著性的YOLOv3船舶目标检测算法研究[J].软件导刊,2020,19(10):146-151. DOI:10.11907/rjdk.201157.
[18] Hu J, Shen L, Albanie S, et al. Squeeze-and-excitation networks[C]//IEEE Transactions on Pattern Analysis and Machine Intelligence. IEEE,:2011-2023. DOI:10.1109/TPAMI.2019.2913372.
[19] Zheng Z H, Wang P, Liu W, et al. Distance-IoU loss:faster and better learning for bounding box regression[J]. Proceedings of the AAAI Conference on Artificial Intelligence, 2020, 34(7):12993-13000. DOI:10.1609/aaai.v34i07.6999.
[20] Hernández G A, König P. Do deep nets really need weight decay and dropout?[EB/OL]. arXiv:1802.07042. (2018-02-21)[2021-12-15]. https://arxiv.org/abs/1802.07042.
[21] Ioffe S, Szegedy C. Batch normalization:accelerating deep network training by reducing internal covariate shift[EB/OL]. arXiv:1502.03167. (2021-03-29)[2021-12-15]. http://arxiv.org/abs/1502.03167v3.
[22] Rezatofighi H, Tsoi N, Gwak J, et al. Generalized intersection over union:a metric and a loss for bounding box regression[C]//2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). June 15-20, 2019, Long Beach, CA, USA. IEEE, 2019:658-666. DOI:10.1109/CVPR.2019.00075.
[23] 谢梦, 刘伟, 杨梦圆, 等. 深度卷积神经网络支持下的遥感影像飞机检测[J]. 测绘通报, 2019(6):19-23. DOI:10.13474/j.cnki.11-2246.2019.0177.
[24] Zhang P, Su W H. Statistical inference on recall, precision and average precision under random selection[C]//20129th International Conference on Fuzzy Systems and Knowledge Discovery. May 29-31, 2012, Chongqing, China. IEEE, 2012:1348-1352. DOI:10.1109/FSKD.2012.6234049.
[25] Bochkovskiy A, Wang C Y, Liao H Y M. YOLOv4:optimal speed and accuracy of object detection[EB/OL]. arXiv:2004.10934. (2020-04-23)[2021-12-15]. https://arxiv.org/abs/2004.10934.
[26] Jocher G, Nishimura K, Mineeva T, et al. YOLOv5[EB/OL]. (2020-08-10)[2021-12-15]. https://github.com/ultralytics/yolov5.
文章导航

/