A Fast Inference Vision Transformer for Automatic Pavement Image Classification and Its Visual Interpretation Method-Reference-Cited by-同舟云学术

A Fast Inference Vision Transformer for Automatic Pavement Image Classification and Its Visual Interpretation Method

Published:2022-04-13 Issue:8 Volume:14 Page:1877
ISSN:2072-4292
Container-title:Remote Sensing
language:en
Short-container-title:Remote Sensing

Author:

Chen Yihan^ORCID,Gu Xingyu,Liu Zhen^ORCID,Liang Jia

Abstract

Traditional automatic pavement distress detection methods using convolutional neural networks (CNNs) require a great deal of time and resources for computing and are poor in terms of interpretability. Therefore, inspired by the successful application of Transformer architecture in natural language processing (NLP) tasks, a novel Transformer method called LeViT was introduced for automatic asphalt pavement image classification. LeViT consists of convolutional layers, transformer stages where Multi-layer Perception (MLP) and multi-head self-attention blocks alternate using the residual connection, and two classifier heads. To conduct the proposed methods, three different sources of pavement image datasets and pre-trained weights based on ImageNet were attained. The performance of the proposed model was compared with six state-of-the-art (SOTA) deep learning models. All of them were trained based on transfer learning strategy. Compared to the tested SOTA methods, LeViT has less than 1/8 of the parameters of the original Vision Transformer (ViT) and 1/2 of ResNet and InceptionNet. Experimental results show that after training for 100 epochs with a 16 batch-size, the proposed method acquired 91.56% accuracy, 91.72% precision, 91.56% recall, and 91.45% F1-score in the Chinese asphalt pavement dataset and 99.17% accuracy, 99.19% precision, 99.17% recall, and 99.17% F1-score in the German asphalt pavement dataset, which is the best performance among all the tested SOTA models. Moreover, it shows superiority in inference speed (86 ms/step), which is approximately 25% of the original ViT method and 80% of some prevailing CNN-based models, including DenseNet, VGG, and ResNet. Overall, the proposed method can achieve competitive performance with fewer computation costs. In addition, a visualization method combining Grad-CAM and Attention Rollout was proposed to analyze the classification results and explore what has been learned in every MLP and attention block of LeViT, which improved the interpretability of the proposed pavement image classification model.

Publisher

MDPI AG

Subject

General Earth and Planetary Sciences

Link

https://www.mdpi.com/2072-4292/14/8/1877/pdf

Reference47 articles.

1. Deep Learning-Based Thermal Image Analysis for Pavement Defect Detection and Classification Considering Complex Pavement Conditions

2. Application of Combining YOLO Models and 3D GPR Images in Road Detection and Maintenance

3. Comparison of deep convolutional neural networks and edge detectors for image-based crack detection in concrete

4. The State-of-the-Art Review on Applications of Intrusive Sensing, Image Processing Techniques, and Machine Learning Methods in Pavement Monitoring and Analysis

5. 3D Visualization of Airport Pavement Quality Based on BIM and WebGL Integration

Cited by 29 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Fusion of transformer attention and CNN features for skin cancer detection;Applied Soft Computing;2024-10

2. Towards automatic phytolith classification using feature extraction and combination strategies;Progress in Artificial Intelligence;2024-07-31

3. Potato leaf disease detection with a novel deep learning model based on depthwise separable convolution and transformer networks;Engineering Applications of Artificial Intelligence;2024-07

4. Automatic extraction and 3D modeling of real road scenes using UAV imagery and deep learning semantic segmentation;International Journal of Digital Earth;2024-06-19

5. AI-based recognition for road distressed from GPR measurement using artificial neural networks;International Conference on Image, Signal Processing, and Pattern Recognition (ISPP 2024);2024-06-13