arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.02470cs.CVcs.AI

基于专用分割模型实现具身智能视觉语言模型(Agentic VLMs)的细粒度车辆损伤评估

Grounding Agentic VLMs with Dedicated Segmentation for Fine-Grained Vehicle Damage Assessment

  • Northeastern University(东北大学)

机构由 AI 辅助整理,请以论文原文为准。

Vishwajeet Shivaji Hogale, Anjali Pai, Nitya Ravi

AI总结:

针对VLMs空间定位不可靠问题,本文提出TinyDamage架构,将空间定位委托给专用多任务分割模型,集成至7节点LangGraph智能体流程,在车辆损伤评估中大幅降低报告虚构率。

AI中文摘要:

视觉语言模型(VLMs)正越来越多地作为推理智能体被部署到实际视觉评估流程中,但其空间定位对于细粒度、视觉模糊的目标仍不可靠。本文针对自动车辆损伤评估场景研究该差距,其中划痕、发丝级裂纹等细粒度缺陷仅占少量像素,梯度信号弱,易与反射和表面纹理混淆。研究发现,当前最先进的VLM(Qwen-VL)在该任务上语义分类准确率达87.3%,但空间层面存在系统性定位问题:会在反射区域虚构损伤、完全遗漏细长划痕,且在要求定位时输出空间不一致。本文提出TinyDamage混合架构,将空间定位委托给专用多任务分割模型,同时让VLM负责语义推理和报告生成。在分割方面,损失函数选择对微小目标定位有显著且未被充分研究的影响:广泛用于类别不平衡的焦点损失会将微小损伤检测降至0,而监督对比目标可显著提升损伤与背景的可分性。本文将该分割模型集成到7节点LangGraph智能体流程中,使VLM每一步生成都基于分割输出,在100份经人工验证报告的受控评估中,该定位将报告虚构率从92%(仅文本)、78%(仅图像)降至31%。本文还引入DET_l,一种宽松的逐类别检测指标,用于评估类别不平衡下的微小目标定位,并报告部署流程的延迟和可靠性特征。

英文摘要:

Vision-language models (VLMs) are increasingly deployed as reasoning agents in real-world visual assessment pipelines, yet their spatial grounding remains unreliable for fine-grained, visually ambiguous targets. We study this gap in the context of automated vehicle damage assessment, where fine-grained defects such as scratches and hairline cracks occupy few pixels, produce weak gradient signal, and are easily confused with reflections and surface texture. We show that a state-of-the-art VLM (Qwen-VL) achieves strong semantic classification accuracy (87.3%) on this task but is systematically ungrounded at the spatial level: it hallucinates damage in reflective regions, misses elongated scratches entirely, and produces spatially inconsistent outputs when prompted for localization. We propose TinyDamage, a hybrid architecture that delegates spatial grounding to a dedicated multi-task segmentation model while reserving the VLM for semantic reasoning and report generation. On the segmentation side, we find that the choice of loss function has an outsized and underexplored effect on tiny-object grounding: focal loss, widely used for class imbalance, collapses tiny-damage detection to zero, while a supervised contrastive objective measurably improves damage/background separability. We integrate the segmentation model into a 7-node LangGraph agent pipeline that grounds every VLM generation step in the segmentation output, and show that this grounding reduces the report hallucination rate from 92% (text-only) and 78% (image-only) to 31% in a controlled evaluation on 100 human-verified reports. We introduce DET_l, a permissive per-category detection metric for evaluating tiny-object grounding under class imbalance, and report latency and reliability characteristics of the deployed pipeline.

补充信息

↑