Volume 26 Issue 6
Jun.  2026
Turn off MathJax
Article Contents
WANG Si-yu, FANG Hong-su, YANG Wei, ZHOU Yong-jun. Road pothole segmentation network based on multi-scale feature fusion and channel feature adaptation[J]. Journal of Traffic and Transportation Engineering, 2026, 26(6): 167-185. doi: 10.19818/j.cnki.1671-1637.2026.031
Citation: WANG Si-yu, FANG Hong-su, YANG Wei, ZHOU Yong-jun. Road pothole segmentation network based on multi-scale feature fusion and channel feature adaptation[J]. Journal of Traffic and Transportation Engineering, 2026, 26(6): 167-185. doi: 10.19818/j.cnki.1671-1637.2026.031

Road pothole segmentation network based on multi-scale feature fusion and channel feature adaptation

doi: 10.19818/j.cnki.1671-1637.2026.031
Funds:

National Key R&D Program of China 2021YFB2601000

Key R&D Program of Shaanxi Province 2024CY2-GJHX-31

More Information
  • Corresponding author: YANG Wei, lecturer, PhD, E-mail: yw@chd.edu.cn
  • Received Date: 2025-01-12
  • Accepted Date: 2025-06-05
  • Rev Recd Date: 2025-03-27
  • Publish Date: 2026-06-28
  • To accurately identify road potholes and the improve model generalization segmentation ability, a road pothole segmentation network (potholes-FBConvNet) based on astrous spatial pyramid fusion (ASPF) and feature adaptation was proposed. The convolutional neural network model (ConvNext) was used as the main structure for multi-scale feature extraction at different network depths. The unified perception and resolution network architecture (UPerNet) was used as the neck fusion network. The multi-level feature representations extracted through the main network were fully utilized, and the ASPF module was introduced into the neck network to further enhance its ability to capture multi-scale contextual feature information. A feature adaptation module was embedded between the neck and decoding head to perform adaptive calibration on the channel and spatial features before neck fusion feature decoding. In this way, the network segmentation performance was effectively improved. A total of 2 097 road pothole segmentation databases were separately collected and meticulously annotated (1 065 generally damaged potholes, 507 completely damaged potholes, 525 severely damaged potholes), and comparative and ablation experiments were conducted in sequence to verify network performance. The experimental results show that, the proposed potholes-FBConvNet model achieves an intersection over union (IoU), similarity coefficient, precision, and recall of 85.43%, 92.14%, 92.96%, and 91.33% on generally damaged potholes. On completely damaged potholes, it reaches 87.76%, 93.48%, 92.33%, and 94.67%. On severely damaged potholes, it reaches 90.76%, 95.15%, 95.41%, and 94.90%. Compared with 16 typical comparison models based on Transformer and Conv, the proposed segmentation network has the optimal generalization segmentation ability and robustness.

     

  • loading
  • [1]
    YAO Chu-xian, CAI Hao-nan, ZHANG Yuan-bo, et al. Road damage detection method based on lightweight vehicle equipment[J]. Bulletin of Surveying and Mapping, 2024(5): 147-150.
    [2]
    CUI Xin-zhuang, HUANG Dan, LIU Lei, et al. A review of mechanics of asphalt pavement disease[J]. Journal of Shandong University (Engineering Science), 2016, 46(5): 68-87.
    [3]
    GUAN Jin-chao, DING Ling, YANG Xu, et al. Pavement surface distress detection in complex scenarios driven by multi-dimensional image fusion[J]. Journal of Traffic and Transportation Engineering, 2024, 24(3): 154-170. doi: 10.19818/j.cnki.1671-1637.2024.03.010
    [4]
    Editorial Department of China Journal of Highway and Transport. Review on China's pavement engineering research: 2024[J]. China Journal of Highway and Transport, 2024, 37(3): 1-81.
    [5]
    Editorial Department of China Journal of Highway and Transport. Review on China's pavement engineering research · 2020[J]. China Journal of Highway and Transport, 2020, 33(10): 1-66.
    [6]
    GAO M X, WANG X, ZHU S L, et al. Detection and segmentation of cement concrete pavement pothole based on image processing technology[J]. Mathematical Problems in Engineering, 2020(1): 1360832.
    [7]
    ABBAS I H, ISMAEL M Q, Automated pavement distress detection using image processing techniques[J]. Engineering, Technology and Applied Science Research, 2021, 11(5): 7702-7708. doi: 10.48084/etasr.4450
    [8]
    KIM S, SEO D, JEON S. Improvement of tiny object segmentation accuracy in aerial images for asphalt pavement pothole detection[J]. Sensors, 2023, 23(13): 5851. doi: 10.3390/s23135851
    [9]
    YE W L, JIANG W, TONG Z, et al. Convolutional neural network for pothole detection in asphalt pavement[J]. Road Materials and Pavement Design, 2021, 22(1): 42-58. doi: 10.1080/14680629.2019.1615533
    [10]
    WANG A D, LANG H, CHEN Z, et al. The two-step method of pavement pothole and raveling detection and segmentation based on deep learning[J]. IEEE Transactions on Intelligent Transportation Systems, 2024, 25(6): 5402-5417. doi: 10.1109/TITS.2023.3340340
    [11]
    LAKMAL H K I S, DISSANAYAKE M B. Pothole detection with image segmentation for advanced driver assisted systems[C]// IEEE. 2020 IEEE International Women in Engineering (WIE) Conference on Electrical and Computer Engineering (WIECON-ECE). New York: IEEE, 2021: 308-311.
    [12]
    YUAN Chao-chun, WANG Jun-xian, HE You-guo, et al. Active obstacle avoidance control of intelligent vehicle based on pothole detection[J]. Journal of Jiangsu University (Natural Science Edition), 2022, 43(5): 504-511, 518.
    [13]
    XUAN Yi-guo, YU Cheng-bo, JIANG Qi-chao, et al. Improved YOLOv7 road crack and pothole detection algorithm[J]. Science Technology and Engineering, 2024, 24(17): 7205-7213.
    [14]
    WANG Xin, LI Qi. Automatic detection of pavement defects based on deep learning[J]. Journal of Optoelectronics· Laser, 2022, 33(11): 1165-1172.
    [15]
    HU Xiao-wei, YAN Yi-xin, WANG Da-wei, et al. Lightweight pavement disease detection based on YOLOM algorithm[J]. China Journal of Highway and Transport, 2024, 37(12): 381-391.
    [16]
    LUO Xiang-long, WANG Yan-bo, PU Ya-ya, et al. Road disease detection RGT-YOLOv7 model under multiple diseases complicated scenarios[J]. Journal of Hunan University (Natural Sciences), 2024, 51(12): 107-118.
    [17]
    GUO Dao-jun, ZHOU Xing-yu, XU Chong-bang, et al. Improved fast tunnel disease segmentation method based on DeepLabV3+ network[J]. Journal of Highway and Transportation Research and Development, 2023, 40(8): 127-135.
    [18]
    REN Yong-jie, WU Li-peng. Municipal road disease recognition based on attention mechanism and novel convolutional neural network[J]. Science Technology and Engineering, 2024, 24(20): 8663-8672.
    [19]
    LI D R, DUAN Z D, HU X Y, et al. Automated classification and detection of multiple pavement distress images based on deep learning[J]. Journal of Traffic and Transportation Engineering (English Edition), 2023, 10(2): 276-290. doi: 10.1016/j.jtte.2021.04.008
    [20]
    GUAN T, CAI J Y, WANG Y, et al. Pavement pothole detection system based on deep learning and binocular vision[J]. Journal of Traffic and Transportation Engineering (English Edition), 2025, 12(4): 1100-1123. doi: 10.1016/j.jtte.2024.08.001
    [21]
    LIU Z, MAO H Z, WU C Y, et al. A ConvNet for the 2020s[C]// IEEE. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). New York: IEEE, 2022: 11966-11976.
    [22]
    VASWANI A, SHAZEER N, PARMAR N, et al. Attention is all you need[PP/OL]. V7. ArXiv (2023-08-02). https://arxiv.org/abs/1706.03762.
    [23]
    HE K M, ZHANG X Y, REN S Q, et al. Deep residual learning for image recognition[C]//IEEE. 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). New York: IEEE, 2016: 770-778.
    [24]
    XIAO TT, LIU Y C, ZHOU B L, et al. Unified perceptual parsing for scene understanding[C]//Springer. Computer Vision-ECCV 2018. Munich: Springer, 2018: 432-448.
    [25]
    CHEN L C, PAPANDREOU G, SCHROFF F, et al. Rethinking atrous convolution for semantic image segmentation[PP/OL]. V3. ArXiv (2017-12-05). https://arxiv.org/abs/1706.05587.
    [26]
    PARK J, WOO S, LEE J Y, et al. BAM: Bottleneck attention module[PP/OL]. V2. ArXiv (2018-07-18). https://arxiv.org/abs/1807.06514.
    [27]
    XIE E Z, WANG W H, YU Z D, et al. SegFormer: Simple and efficient design for semantic segmentation with Transformers[PP/OL]. V3. ArXiv (2021-10-28). https://arxiv.org/abs/2105.15203.
    [28]
    CHU XX, TIAN Z, WANG Y Q, et al. Twins: Revisiting the design of spatial attention in vision transformers[PP/OL]. V4. ArXiv (2021-09-30). https://arxiv.org/abs/2104.13840.
    [29]
    LIU Z, LIN Y T, CAO Y, et al. Swin Transformer: Hierarchical vision transformer using shifted windows[C]// IEEE. 2021 IEEE/CVF International Conference on Computer Vision (ICCV). New York: IEEE, 2022: 9992-10002.
    [30]
    RANFTL R, BOCHKOVSKIY A, KOLTUN V. Vision Transformers for dense prediction[C]//IEEE. 2021 IEEE/CVF International Conference on Computer Vision (ICCV). New York: IEEE, 2022: 12159-12168.
    [31]
    DOSOVITSKIY A, BEYER L, KOLESNIKOV A, et al. An image is worth 16×16 words: Transformers for image recognition at scale[PP/OL]. V2. ArXiv (2021-06-03). https://arxiv.org/abs/2010.11929.
    [32]
    HOWARD A, SANDLER M, CHEN B, et al. Searching for MobileNetV3[C]//IEEE. 2019 IEEE/CVF International Conference on Computer Vision (ICCV). New York: IEEE, 2019: 1314-1324.
    [33]
    POUDEL R P K, LIWICKI S, CIPOLLA R. Fast-SCNN: Fast semantic segmentation network[PP/OL]. V1. ArXiv (2019-02-12). https://arxiv.org/abs/1902.04502.
    [34]
    FAN M Y, LAI S Q, HUANG J S, et al. Rethinking BiSeNet for real-time semantic segmentation[C]//IEEE. 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). New York: IEEE, 2021: 9711-9720.
    [35]
    YU W H, LUO M, ZHOU P, et al. MetaFormer is actually what you need for vision[C]//IEEE. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). New York: IEEE, 2022: 10809-10819.
    [36]
    RONNEBERGER O, FISCHER P, BROX T. U-Net: Convolutional networks for biomedical image segmentation[C]// Springer. Medical Image Computing and Computer-assisted Intervention-MICCAI 2015. Munich: Springer, 2015: 234-241.
    [37]
    YU C Q, GAO C X, WANG J B, et al. BiSeNet V2: Bilateral network with guided aggregation for real-time semantic segmentation[J]. International Journal of Computer Vision, 2021, 129(11): 3051-3068. doi: 10.1007/s11263-021-01515-2
    [38]
    CHEN L C, ZHU Y K, PAPANDREOU G, et al. Encoder-decoder with atrous separable convolution for semantic image segmentation[C]//Springer. Computer Vision-ECCV 2018. Munich: Springer, 2018: 833-851.
    [39]
    KIRILLOV A, WU Y X, HE K M, et al. PointRend: Image segmentation as rendering[C]//IEEE. 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). New York: IEEE, 2020: 9796-9805.
    [40]
    ZHANG W W, PANG J M, CHEN K, et al. K-net: Towards unified image segmentation[PP/OL]. V2. ArXiv (2021-11-01). https://arxiv.org/abs/2106.14855.
    [41]
    JIN Z C, LIU B, CHU Q, et al. ISNet: Integrate image-level and semantic-level context for semantic segmentation[C]// IEEE. 2021 IEEE/CVF International Conference on Computer Vision (ICCV). New York: IEEE, 2022: 7169-7178.
    [42]
    ZHU Z, XU M D, BAI S, et al. Asymmetric non-local neural networks for semantic segmentation[C]//IEEE. 2019 IEEE/CVF International Conference on Computer Vision (ICCV). New York: IEEE, 2019: 593-602.
    [43]
    ZHU Yan-jie, WANG Yu-chen, XIONG Wen, et al. Few-shot model for extracting inspection report information based on bridge inspection domain-task transfer[J]. Journal of Traffic and Transportation Engineering, 2025, 25(1): 248-262.

Catalog

    Article Metrics

    Article views (43) PDF downloads(7) Cited by()
    Related

    /

    DownLoad:  Full-Size Img  PowerPoint
    Return
    Return