Volume 26 Issue 6
Jun.  2026
Turn off MathJax
Article Contents
YANG Yang, CHEN Xian-tian, WANG Jian-yu, PU Zi-yuan, ZHAO Hong-zhuan, YUAN Zhen-zhou. Construction of VSSM-CNN detection network for nighttime road traffic accidents[J]. Journal of Traffic and Transportation Engineering, 2026, 26(6): 137-152. doi: 10.19818/j.cnki.1671-1637.2026.076
Citation: YANG Yang, CHEN Xian-tian, WANG Jian-yu, PU Zi-yuan, ZHAO Hong-zhuan, YUAN Zhen-zhou. Construction of VSSM-CNN detection network for nighttime road traffic accidents[J]. Journal of Traffic and Transportation Engineering, 2026, 26(6): 137-152. doi: 10.19818/j.cnki.1671-1637.2026.076

Construction of VSSM-CNN detection network for nighttime road traffic accidents

doi: 10.19818/j.cnki.1671-1637.2026.076
Funds:

National Natural Science Foundation of China 52572336

Natural Science Foundation of Beijing E2024210149

Cultivation Project Funds for Beijing University of Civil Engineering and Architecture X25034

Fundamental Research Funds for the Central Universities 2024JBRC009

More Information
  • Corresponding author: WANG Jian-yu, associate professor, PhD, E-mail: wangjianyu@bucea.edu.cn
  • Received Date: 2025-08-09
  • Accepted Date: 2025-09-28
  • Rev Recd Date: 2025-09-17
  • Publish Date: 2026-06-28
  • To enhance the automatic detection performance of road traffic accidents in nighttime scenarios, a VSSM-CNN encoder-decoder architecture was constructed based on a visual state space model (VSSM) and a convolutional neural network (CNN), and an unsupervised traffic accident detection framework oriented to this scenario was proposed. By referencing the idea of feature fusion and based on existing visible-light images as appearance features, fine-grained optical flow information of video sequences was further extracted and processed using a recurrent all-pairs field transform (RAFT) optical flow estimation algorithm combined with a convolutional long short-term memory (ConvLSTM) module to represent traffic motion states. The two types of features, appearance and motion, were fused and input into a feature encoder with the VSSM as a backbone network for global feature extraction, and a CNN was adopted as a decoder architecture to recover images layer by layer to enhance local details, so as to strengthen feature extraction efficiency and utilization effects. A triple loss function combination of mean squared error, mean absolute error, and structural similarity was adopted, and weight proportions were determined through Bayesian optimization to improve the robustness of the model to abnormal structures and noise during nighttime image reconstruction. Research results indicate that the area under the receiver operating characteristic curve (ROC-AUC) and area under the precision-recall curve (PR-AUC) of the proposed method are 0.818 and 0.765, respectively, which improve by 22.6% and 20.3%, respectively, compared with traditional generative networks; among them, the ROC-AUC achieves the highest value in the comparison with multiple existing methods; the triple loss function shows the strongest model performance improvement ability in ablation experiments; the research method achieves the lowest model complexity and a favorable detection speed, fully demonstrating its stability in anomaly recognition and potential for deployment and application. The research results effectively improve the traffic accident detection ability in nighttime scenarios, can be further extended to traffic accident detection in low-illumination and low-visibility scenarios, and provide a feasible technical solution and reference for related applications.

     

  • loading
  • [1]
    WANG Chen, ZHOU Wei, YAN Jun-yi, et al. Improved two-stream network for vision-based traffic accident detection[J]. China Journal of Highway and Transport, 2023, 36(5): 185-196.
    [2]
    YANG Yang, WANG Wen-hui, WU Xian-yu, et al. Review of the research toward freeway unconventional traffic accidents[J]. Journal of Basic Science and Engineering, 2024, 32(3): 601-626.
    [3]
    WANG W C, YANG Y, YANG X B, et al. A negative binomial Lindley approach considering spatiotemporal effects for modeling traffic crash frequency with excess zeros[J]. Accident Analysis & Prevention, 2024, 207: 107741.
    [4]
    HUANG Si-de, HUANG Rong-jun, ZHANG Yan-kun, et al. An algorithm for automatic traffic incident detection based on surveillance video[J]. Journal of Highway and Transportation Research and Development, 2024, 41(1): 169-176.
    [5]
    GUO Yan-yong, LIU Pei, YUAN Quan, et al. Review on research of road traffic safety of connected and automated vehicles[J]. Journal of Traffic and Transportation Engineering, 2023, 23(5): 19-38. doi: 10.19818/j.cnki.1671-1637.2023.05.002
    [6]
    General Office of the Ministry of Public Security, General Office of the National Health Commission. Notice on improving the long-term mechanism of police-medical joint rescue and treatment for road traffic accidents: Gong Jiao Guan 〔2020〕 No. 161[EB/OL]. (2020-07-02)[2025-08-04]. http://www.nhc.gov.cn/vzvgj/s3594q/202007/c41e38d7db744ad9890ecb5d710ab59b.shtml.
    [7]
    YAN Ying, WANG Yu-ying, ZHOU Xuan, et al. Analysis of autonomous vehicle accident severity factors based on hybrid model[J]. Journal of Traffic and Transportation Engineering, 2025, 25(1): 184-196. doi: 10.19818/j.cnki.1671-1637.2025.01.013
    [8]
    YANG Yang, HE Kun, WANG Yun-peng, et al. Spatial transplantation for modeling of freeway traffic crash risk based on dynamic traffic flow[J]. Journal of Transportation Systems Engineering and Information Technology, 2023, 23(3): 174-186.
    [9]
    NIU Shi-feng, MU Jun-jie, PU Ze-yu, et al. Hierarchical assessment method for vehicle accident risk classification based on temporal misalignment causal variables and sub-objective training[J]. Journal of Traffic and Transportation Engineering, 2025, 25(6): 293-306. doi: 10.19818/j.cnki.1671-1637.2025.06.024
    [10]
    YANG Yang, YUAN Zhen-zhou, WANG Yin-hai, et al. Freeway crash risk identification based on a new improved method of WOMDI-Apriori algorithm[J]. Journal of Transportation Engineering, 2021, 21(6): 1-10, 16.
    [11]
    BATANINA E, BEKKOUCH I E I, YOUSSRY Y, et al. Domain adaptation for car accident detection in videos[C]//IEEE. 2019 Ninth International Conference on Image Processing Theory, Tools and Applications (IPTA). New York: IEEE, 2019: 1-6.
    [12]
    FANG J W, QIAO J H, BAI J, et al. Traffic accident detection via self-supervised consistency learning in driving scenarios[J]. IEEE Transactions on Intelligent Transportation Systems, 2022, 23(7): 9601-9614. doi: 10.1109/TITS.2022.3157254
    [13]
    PASHAEI A, GHATEE M, SAJEDI H. Convolution neural network joint with mixture of extreme learning machines for feature extraction and classification of accident images[J]. Journal of Real-time Image Processing, 2020, 17(4): 1051-1066. doi: 10.1007/s11554-019-00852-3
    [14]
    IJJINA E P, CHAND D, GUPTA S, et al. Computer vision-based accident detection in traffic surveillance[C]//IEEE. 2019 10th International Conference on Computing, Communication and Networking Technologies (ICCCNT). New York: IEEE, 2019: 1-6.
    [15]
    JIANG Y, WANG Y H, ZHAO M H, et al. Nighttime traffic object detection via adaptively integrating event and frame domains[J]. Fundamental Research, 2025, 5(4): 1633-1644. doi: 10.1016/j.fmre.2023.08.004
    [16]
    YUE Y G, ZHANG S, WU Z Y, et al. Research on nighttime road visibility monitoring based on video images[J]. Traffic Injury Prevention, 2026, 27(3): 287-295. doi: 10.1080/15389588.2025.2495203
    [17]
    HOSSAIN A, SUN X D, SHAHRIER M, et al. Exploring nighttime pedestrian crash patterns at intersection and segments: Findings from the machine learning algorithm[J]. Journal of Safety Research, 2023, 87: 382-394. doi: 10.1016/j.jsr.2023.08.010
    [18]
    REN J S, WANG W, WANG J W, et al. An unsupervised feature learning approach to improve automatic incident detection[C]//IEEE. 2012 15th International IEEE Conference on Intelligent Transportation Systems. New York: IEEE, 2012: 172-177.
    [19]
    YAO Y, XU M Z, WANG Y C, et al. Unsupervised traffic accident detection in first-person videos[C]//IEEE. 2019 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). New York: IEEE, 2019: 273-280.
    [20]
    SINGH D, MOHAN C K. Deep spatio-temporal representation for detection of road accidents using stacked autoencoder[J]. IEEE Transactions on Intelligent Transportation Systems, 2019, 20(3): 879-887. doi: 10.1109/TITS.2018.2835308
    [21]
    ZHOU W, WEN L H, ZHAN Y F, et al. An appearance-motion network for vision-based crash detection: Improving the accuracy in congested traffic[J]. IEEE Transactions on Intelligent Transportation Systems, 2023, 24(12): 13742-13755. doi: 10.1109/TITS.2023.3297589
    [22]
    CHEN P, ZHANG W W, XIAO Z Y, et al. Traffic accident detection based on deformable frustum proposal and adaptive space segmentation[J]. Computer Modeling in Engineering & Sciences, 2022, 130(1): 97-109.
    [23]
    ZHANG Qi, ZHOU Wei, HU Wei-chao, et al. Accident detection method integrating progressive domain adaptation and cross-attention[J]. Computer Engineering and Applications, 2025, 61(6): 349-360.
    [24]
    TRAN T M, BUI D C, NGUYEN T V, et al. Transformer-based spatio-temporal unsupervised traffic anomaly detection in aerial videos[J]. IEEE Transactions on Circuits and Systems for Video Technology, 2024, 34(9): 8292-8309. doi: 10.1109/TCSVT.2024.3376399
    [25]
    LIANG R Q, LI Y M, YI Y X, et al. A memory-augmented multi-task collaborative framework for unsupervised traffic anomaly detection in driving videos[J]. Pattern Recognition, 2025, 168: 111789. doi: 10.1016/j.patcog.2025.111789
    [26]
    LUO W X, LIU W, LIAN D Z, et al. Video anomaly detection with sparse coding inspired deep neural networks[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2021, 43(3): 1070-1084. doi: 10.1109/TPAMI.2019.2944377
    [27]
    ZHANG Y, NIE X S, HE R D, et al. Normality learning in multispace for video anomaly detection[J]. IEEE Transactions on Circuits and Systems for Video Technology, 2021, 31(9): 3694-3706. doi: 10.1109/TCSVT.2020.3039798
    [28]
    QIE K, WANG J Y, LI Z H, et al. Recognition of occluded pedestrians from the driver's perspective for extending sight distance and ensuring driving safety at signal-free intersections[J]. Digital Transportation and Safety, 2024, 3(2): 65-74. doi: 10.48130/dts-0024-0007
    [29]
    SHI Yang-yu, XIE Cheng-jie, ZHENG Di-wen, et al. Multi-scale anomaly behavior detection method based on Mamba-CNN[J/OL]. Journal of Beijing University of Aeronautics and Astronautics, (2024-11-15). https://doi.org/10.13700/j.bh.1001-5965.2024.0416.
    [30]
    PENG Z L, GUO Z H, HUANG W, et al. Conformer: Local features coupling global representations for recognition and detection[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2023, 45(8): 9454-9468. doi: 10.1109/TPAMI.2023.3243048
    [31]
    SRINIVAS A, LIN T Y, PARMAR N, et al. Bottleneck transformers for visual recognition[C]//IEEE. 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). New York: IEEE, 2021: 16514-16524.
    [32]
    TEED Z, DENG J. RAFT: Recurrent all-pairs field transforms for optical flow[C]//ECCV. Computer Vision-ECCV 2020. Berlin: Springer International Publishing, 2020: 402-419.
    [33]
    JIAO J B, LIU Y, LIU Y F, et al. VMamba: Visual state space model[C]//NeurIPS. Advances in Neural Information Processing Systems 37. Vancouver: Curran Neural Information Processing Systems Foundation, Inc. (NeurIPS), 2024: 103031-103063.
    [34]
    DOSOVITSKIY A, FISCHER P, ILG E, et al. FlowNet: Learning optical flow with convolutional networks[C]//IEEE. 2015 IEEE International Conference on Computer Vision (ICCV). New York: IEEE, 2015: 2758-2766.
    [35]
    SUN D Q, YANG X D, LIU M Y, et al. PWC-net: CNNs for optical flow using pyramid, warping, and cost volume[C]// IEEE. 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). New York: IEEE, 2018: 8934-8943.
    [36]
    CHEN H J, WANG Z, QIN H D, et al. CFDHI-net: Correlation-driven feature decoupling and hierarchical integration network for RGB-thermal semantic segmentation[J]. IEEE Transactions on Intelligent Transportation Systems, 2025, 26(10): 17173-17184. doi: 10.1109/TITS.2025.3581609
    [37]
    FAN J, CHEN M, GU Z Y, et al. SSIM over MSE: A new perspective for video anomaly detection[J]. Neural Networks, 2025, 185: 107115. doi: 10.1016/j.neunet.2024.107115
    [38]
    ISOLA P, ZHU J Y, ZHOU T H, et al. Image-to-image translation with conditional adversarial networks[C]//IEEE. 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). New York: IEEE, 2017: 1125-1134.
    [39]
    KLEIN A, FALKNER S, BARTELS S, et al. Fast Bayesian optimization of machine learning hyperparameters on large datasets[C]//PMLR. Proceedings of the 20th International Conference on Artificial Intelligence and Statistics. Fort Lauderdale: PMLR, 2017: 528-536.
    [40]
    WU J, TOSCANO-PALMERIN S, FRAZIER P I, et al. Practical multi-fidelity Bayesian optimization for hyperparameter tuning[C]//PMLR. Proceedings of the 35th Uncertainty in Artificial Intelligence Conference. Fort Lauderdale: PMLR, 2020: 788-798.
    [41]
    SHAH A P, LAMARE J B, NGUYEN-ANH T, et al. CADP: A novel dataset for CCTV traffic camera based accident analysis[C]//IEEE. 2018 15th IEEE International Conference on Advanced Video and Signal Based Surveillance (AVSS). New York: IEEE, 2018: 1-9.
    [42]
    HUANG X H, HE P, RANGARAJAN A, et al. Intelligent intersection: Two-stream convolutional networks for real-time near-accident detection in traffic video[J]. ACM Transactions on Spatial Algorithms and Systems, 2020, 6(2): 1-28.
    [43]
    HASAN M, CHOI J, NEUMANN J, et al. Learning temporal regularity in video sequences[C]//IEEE. 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). New York: IEEE, 2016: 733-742.
    [44]
    TRAN D, BOURDEV L, FERGUS R, et al. Learning spatiotemporal features with 3D convolutional networks[C]// IEEE. 2015 IEEE International Conference on Computer Vision (ICCV). New York: IEEE, 2015: 4489-4497.
    [45]
    GONG D, LIU L Q, LE V, et al. Memorizing normality to detect anomaly: Memory-augmented deep autoencoder for unsupervised anomaly detection[C]//IEEE. 2019 IEEE/CVF International Conference on Computer Vision (ICCV). New York: IEEE, 2019: 1705-1714.
    [46]
    PARK H, NOH J, HAM B. Learning memory-guided normality for anomaly detection[C]//IEEE. 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). New York: IEEE, 2020: 14372-14381.
    [47]
    DOSHI K, YILMAZ Y. Online anomaly detection in surveillance videos with asymptotic bounds on false alarm rate[C]//Springer International Publishing. Pattern Recognition-ICPR International Workshops and Challenges. Berlin: Springer International Publishing, 2021: 292-304.

Catalog

    Article Metrics

    Article views (71) PDF downloads(17) Cited by()
    Related

    /

    DownLoad:  Full-Size Img  PowerPoint
    Return
    Return