arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于视觉Transformer的多级特征融合用于多标签下水道缺陷分类

Vision Transformer-Based Multi-Level Feature Fusion for Multi-Label Sewer Defect Classification

Xu Fang, Zhuoran Wang, Qing Li, Shengyu Zhang, Guanzhi Deng, Jianbiao He, Qingquan Li

arXiv 2609.11375首次发表:更新:

发表机构

Shenzhen Polytechnic University; City University of Hong Kong; Pengcheng Laboratory; Shenzhen University(深圳职业技术大学; 香港城市大学; 鹏城实验室; 深圳大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对下水道缺陷多标签分类中精度与计算复杂度的平衡难题,提出分层视觉Transformer及两种轻量级架构,在Sewer-ML上排名第一并大幅降低参数。

AI 中文摘要

下水道缺陷的自动化分类对于基础设施状况评估和维护决策至关重要,但现有的深度学习方法在大规模多标签场景中难以平衡分类准确性和计算复杂度。本研究开发了Sewer-Transformer-ML,一种具有多级特征融合的分层视觉Transformer,以及两种轻量级架构Sewer-MobileNet-ML和Sewer-Mobile-TransNet,用于资源受限的检测场景。在Sewer-ML测试集上,Sewer-Transformer-ML-Base实现了65.68%的$F2_{\ ext{CIW}}$和92.68%的$F1_{\ ext{Normal}}$,在公开排行榜上排名第一,在$F2_{\ ext{CIW}}$上比排名第二的方法高出7.6个百分点。Sewer-MobileNet-ML仅用17M参数就实现了65.73%的$F2_{\ ext{CIW}}$,相对于基础模型参数减少了约95%。在标准Sewer-Capsule数据划分下,Sewer-Mobile-TransNet实现了96.43%的分类准确率。当训练集缩减到1,177张图像时,在Sewer-ML上的预训练持续提升了模型性能。消融实验进一步表明,直接拼接对Transformer特征更有效,而基于注意力的融合更好地支持多尺度CNN特征。这些发现为自动化下水道检测、轻量级模型设计以及跨土木基础设施检测平台的适应性提供了计算基础。

英文摘要

Automated classification of sewer defects is essential for infrastructure condition assessment and maintenance decision-making, but existing deep learning methods struggle to balance classification accuracy and computational complexity in large-scale multi-label scenarios. This study develops Sewer-Transformer-ML, a hierarchical vision Transformer with multi-level feature fusion, together with two lightweight architectures, Sewer-MobileNet-ML and Sewer-Mobile-TransNet, for resource-constrained inspection scenarios. On the Sewer-ML test set, Sewer-Transformer-ML-Base achieved an $F2_{\text{CIW}}$ of 65.68% and an $F1_{\text{Normal}}$ of 92.68%, ranking first on the public leaderboard and exceeding the second-ranked method by 7.6 percentage points in $F2_{\text{CIW}}$. Sewer-MobileNet-ML achieved an $F2_{\text{CIW}}$ of 65.73% with only 17 M parameters, representing an approximately 95% parameter reduction relative to the base model. Under the standard Sewer-Capsule data split, Sewer-Mobile-TransNet achieved 96.43% classification accuracy. When the training set was reduced to 1,177 images, pretraining on Sewer-ML consistently improved model performance. Ablation experiments further showed that direct concatenation was more effective for Transformer features, whereas attention-based fusion better supported multiscale CNN features. These findings provide a computational basis for automated sewer inspection, lightweight model design, and adaptation across civil infrastructure inspection platforms.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑