arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

WADE:面向紧凑视觉-语言模型的多实例漂浮垃圾定位推理标注基准

WADE: A Reasoning-Annotated Benchmark for Multi-Instance Floating-Waste Grounding with Compact Vision-Language Models

Md. Asaduzzaman Shuvo, Ahsan Farabi, Md. Abdul Ahad Minhaz, Mahedi Hasan, Israt Khandaker, Ibrahim Khalil Shanto, Muhammad Nomani Kabir

arXiv 2608.22950首次发表:更新:

发表机构

United International University(联合国际大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对内陆漂浮垃圾监测需求,研究推出WADE基准并评估6种VLMs,微调Qwen3-VL-2B可提升性能但仍有大量实例未被检测,为紧凑VLMs的漂浮垃圾定位提供挑战基准。

AI 中文摘要

内陆水道中的漂浮垃圾威胁水生生态系统,需要在杂乱多目标条件下进行及时监测。现有水生垃圾数据集的地理覆盖范围有限,多实例标注稀疏,除边界框和标签外几乎没有其他监督信息,因此紧凑视觉-语言模型(VLMs)在联合定位、分类、计数和解释漂浮垃圾方面的评估不足。我们推出WADE,这是一个带有推理标注的基准,包含来自孟加拉国农村的2167张图像、13608个边界框以及10个垃圾类别。每个标注都与涵盖视觉线索、可能混淆项和判别特征的类别级识别规则相关联。我们使用检测、计数和幻觉指标,在零样本、少样本、推理引导和微调设置下评估6种VLMs。为实现资源高效适配,我们使用QLoRA联合微调Qwen3-VL-2B,使其适应边界框、标签和推理链。微调后,召回率从0.0248提升至0.2339,F1值从0.0257提升至0.2163,同时图像级幻觉从0.6836降至0.0883。然而,超过四分之三的实例仍未被检测到,这表明WADE是针对紧凑VLMs进行密集漂浮垃圾定位的具有挑战性的基准。

英文摘要

Floating waste in inland waterways threatens aquatic ecosystems and requires timely monitoring under cluttered, multi-object conditions. Existing aquatic-waste datasets provide limited geographic coverage, sparse multi-instance annotations, and little supervision beyond boxes and labels. Compact vision-language models (VLMs) therefore remain insufficiently evaluated for jointly localizing, classifying, counting, and explaining floating waste. We introduce WADE, a reasoning-annotated benchmark containing 2,167 images from rural Bangladesh, 13,608 bounding boxes, and ten waste categories. Each annotation is associated with class-level recognition rules covering visual cues, likely confusions, and discriminative features. We evaluate six VLMs under zero-shot, two-shot, reasoning-guided, and fine-tuned settings using detection, counting, and hallucination metrics. For resource-efficient adaptation, we jointly fine-tune Qwen3-VL-2B on boxes, labels, and reasoning chains using QLoRA. Fine-tuning increases recall from 0.0248 to 0.2339 and F1 from 0.0257 to 0.2163, while reducing image-level hallucination from 0.6836 to 0.0883. However, over three-quarters of instances remain undetected, establishing WADE as a challenging benchmark for dense floating-waste grounding with compact VLMs.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑