Skip to main navigation Skip to search Skip to main content

VTECSeg: Edge-aware hybrid CNN-vision transformer network with zero-shot vision-language-guided region proposals for crack segmentation

Research output: Contribution to journalArticlepeer-review

Abstract

While vision-based crack segmentation automates construction inspection, leading to enhanced safety and long-term cost reduction, it is still challenged by complex crack geometries, cluttered environments, and computational constraints. This paper proposes VTECSeg, an encoder-decoder deep learning (DL) architecture combining a Vision Transformer backbone with a Convolutional Neural Network-based edge encoder to enable local crack detail preservation and global contextual reasoning. To further improve crack segmentation accuracy, a zero-shot Vision-Language Model-guided region proposal mechanism is introduced. Results demonstrate that VTECSeg outperformed several baseline architectures, achieving a precision of 87.32%, a recall of 84.21%, an F1-score of 85.74%, a mean Intersection-over-Union (mIoU) of 86.15%, and an Intersection-over-Union (IoU) of 79.5% on the DeepCrack dataset. This paper contributes to the body of knowledge by developing an edge-aware hybrid crack segmentation framework combined with a zero-shot vision-language-guided region proposal stage for high-fidelity and improved crack segmentation, thus advancing vision-based infrastructure inspection and automation.

Original languageEnglish (US)
Article number106998
JournalAutomation in Construction
Volume188
DOIs
StatePublished - Aug 2026

All Science Journal Classification (ASJC) codes

  • Control and Systems Engineering
  • Civil and Structural Engineering
  • Building and Construction

Keywords

  • Computer vision
  • Convolutional neural networks
  • Crack segmentation
  • Deep learning
  • Inspection automation
  • Region proposals
  • Vision transformers
  • Vision-language models

Fingerprint

Dive into the research topics of 'VTECSeg: Edge-aware hybrid CNN-vision transformer network with zero-shot vision-language-guided region proposals for crack segmentation'. Together they form a unique fingerprint.

Cite this