arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.05859cs.CV

AVA-VLM:用于野外建筑工地监测的自适应视觉注意力视觉语言模型

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring

  • University of California, Los Angeles(加利福尼亚大学洛杉矶分校)
  • AI+KCIR Global Resilience Research Center(人工智能与韩国气候变化影响韧性研究中心)
  • SmartInside AI Co., Ltd.(思玛特因赛德人工智能有限公司)
  • Sungkyunkwan University(成均馆大学)

机构由 AI 辅助整理,请以论文原文为准。

Younggun Kim, Taeheon Kim, Youngseo Kim, Seunghee Park

AI总结:

针对野外建筑工地监测,提出AVA-VLM自适应视觉注意力视觉语言模型,遵循粗到细推理策略,先对低分辨率全局图像推理,按需请求高分辨率局部裁剪,并引入区域感知思维链数据集,提升了低分辨率下的可靠性并减少视觉令牌使用。

AI中文摘要:

视觉语言模型(VLMs)在建筑工地监测方面很有前景,近期针对建筑定制的VLMs主要通过从单个全局图像进行直接问答式微调来适应预训练的VLMs。我们认为这种直接范式在操作范围、低分辨率输入下的可靠性和推理效率方面,对于野外部署仍然有限。为应对这些挑战,我们提出了AVA-VLM,一种遵循人类启发的粗到细推理策略的自适应视觉注意力视觉语言模型。AVA-VLM首先对低分辨率全局图像进行推理,仅在需要详细检查时选择性地请求高分辨率局部裁剪,类似于人类检查员放大难以看清但重要的区域。我们还引入了一个区域感知思维链数据集,教导模型何时检查、在哪里裁剪以及如何使用局部证据。实验表明,AVA-VLM在长距离和低分辨率条件下提高了可靠性,同时大幅减少了视觉令牌的使用。

英文摘要:

Existing construction-site Vision-Language Model (VLM) studies have primarily adapted pretrained VLMs through direct QA-style fine-tuning from a single global image, but we argue that this paradigm remains limited in operational range, reliability under reduced-resolution inputs, and inference efficiency. To address these limitations, we propose AVA-VLM, an Adaptive Visual Attention-Vision Language Model that follows a human-inspired coarse-to-fine strategy: it first reasons over a low-resolution global image and requests a high-resolution local crop only when detailed inspection is needed. We further introduce a region-aware Chain-of-Thought dataset that teaches when to inspect, where to crop, and how to use local evidence. Experiments show that, for violation identification, AVA-VLM improves overall F1 from 62.0 to 75.1 while using only 30.6% of the baseline visual-token budget; for long-distance PPE-violation cases, F1 improves from 16.0 to 63.6. These results demonstrate AVA-VLM's improved robustness to distant and reduced-resolution visual evidence with substantially lower visual-token usage.

↑