arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.27457cs.CV

超越平衡准确率:面向无人机电力线巡检中视觉-语言与纯视觉缺陷评估的分辨率与奇偶性受控基准

Beyond Balanced Accuracy: A Resolution and Parity-Controlled Benchmark for Vision-Language and Vision-Only Defect Assessment in UAV Power-Line Inspection

Linghao Zhang, Siyu Xiang, Junwei Kuang, Peiyu Yi

首次发表
浏览论文内容

中文总结 AI 辅助

本研究构建ElecVQA-Bench基准,系统审计VLM与纯视觉模型在无人机电力线缺陷评估中的性能,发现VLM优势依赖评估设置,经严格受控实验后并不稳健,强调基准审计而非普遍优越性。

中文摘要 AI 辅助

视觉-语言模型(VLM)常被报道在无人机(UAV)电力线缺陷评估任务中优于特定任务的视觉骨干网络。我们在ElecVQA-Bench上检验了这一论断,该基准包含56,972个样本,源自公开的InsPLAD数据集,并涵盖六种评估选择:划分方式、评估样本集、标签空间、重复实验、输入分辨率及辅助信息。在匹配划分下,Swin Transformer与最强的适配VLM在二元筛查中仅相差0.03分。在七类缺陷分类中,将视觉骨干网络从224像素提升至VLM预处理器的实测像素预算后,与InternVL3.5-8B的差距由+20.53收窄至-0.57分(ResNet-50),由+23.67收窄至+4.70分(Swin-T)。像素预算审计使Qwen3-VL-8B的宏召回率变化了10.78分,然而源像素匹配的InternVL对照组仍使Qwen领先7.43至13.61分,同时使用的视觉令牌减少56%,因此源像素或令牌预算均无法解释两个VLM之间的差异。两次种子的全局重复实验使Qwen的二元准确率和七类宏召回率分别变化了0.86和1.02分。在按划分重训练后,Qwen在裁剪或图像级别均不领先,且14个骨干网络、三种种子的重复实验在不同种子间翻转了符号,Qwen的平均共同六类宏召回率为0.9085,而ResNet-50为0.9509。没有任何划分机制能产生经多重比较校正后仍存活的家族级优势。本研究支持基准审计的贡献,而非VLM普遍优越的笼统论断。

英文摘要

Vision-language models (VLMs) are often reported to outperform task-specific vision backbones for unmanned aerial vehicle (UAV) power-line defect assessment. We test that claim on ElecVQA-Bench, a 56,972-item benchmark derived from the public InsPLAD dataset, across six evaluation choices: partition, evaluated item set, label space, replication, input resolution, and side information. On a matched partition, a Swin Transformer and the strongest adapted VLM differ by only 0.03 points at binary screening. At seven-way defect typing, increasing the vision backbones from 224 px to the measured pixel budget of the VLM preprocessor narrows the gap against InternVL3.5-8B from +20.53 to -0.57 points for ResNet-50 and from +23.67 to +4.70 points for Swin-T. A pixel-budget audit shifts Qwen3-VL-8B macro recall by 10.78 points, yet a source-pixel-matched InternVL control still leaves Qwen ahead by 7.43 to 13.61 points while using 56% fewer visual tokens, so neither source pixels nor token budget explains the difference between the two VLMs. A two-seed global replication changes Qwen binary accuracy and seven-way macro recall by 0.86 and 1.02 points. After split-specific retraining, Qwen does not lead at crop or image level, and a 14-tower, three-seed replication reverses the sign across seeds, giving mean common-six macro recall of 0.9085 for Qwen against 0.9509 for ResNet-50. No split regime yields a family-level advantage that survives multiple-comparison correction. The study supports a benchmark-audit contribution rather than a general claim of VLM superiority.

发表机构

  • State Grid Sichuan Electric Power Research Institute(国网四川省电力公司电力科学研究院)
  • Power System Security and Operation Key Laboratory of Sichuan Province(四川省电力系统安全与运行重点实验室)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑