发表机构
United International University(联合国际大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对内陆漂浮垃圾监测需求,研究推出WADE基准并评估6种VLMs,微调Qwen3-VL-2B可提升性能但仍有大量实例未被检测,为紧凑VLMs的漂浮垃圾定位提供挑战基准。
AI 中文摘要
内陆水道中的漂浮垃圾威胁水生生态系统,需要在杂乱多目标条件下进行及时监测。现有水生垃圾数据集的地理覆盖范围有限,多实例标注稀疏,除边界框和标签外几乎没有其他监督信息,因此紧凑视觉-语言模型(VLMs)在联合定位、分类、计数和解释漂浮垃圾方面的评估不足。我们推出WADE,这是一个带有推理标注的基准,包含来自孟加拉国农村的2167张图像、13608个边界框以及10个垃圾类别。每个标注都与涵盖视觉线索、可能混淆项和判别特征的类别级识别规则相关联。我们使用检测、计数和幻觉指标,在零样本、少样本、推理引导和微调设置下评估6种VLMs。为实现资源高效适配,我们使用QLoRA联合微调Qwen3-VL-2B,使其适应边界框、标签和推理链。微调后,召回率从0.0248提升至0.2339,F1值从0.0257提升至0.2163,同时图像级幻觉从0.6836降至0.0883。然而,超过四分之三的实例仍未被检测到,这表明WADE是针对紧凑VLMs进行密集漂浮垃圾定位的具有挑战性的基准。
英文摘要
Floating waste in inland waterways threatens aquatic ecosystems and requires timely monitoring under cluttered, multi-object conditions. Existing aquatic-waste datasets provide limited geographic coverage, sparse multi-instance annotations, and little supervision beyond boxes and labels. Compact vision-language models (VLMs) therefore remain insufficiently evaluated for jointly localizing, classifying, counting, and explaining floating waste. We introduce WADE, a reasoning-annotated benchmark containing 2,167 images from rural Bangladesh, 13,608 bounding boxes, and ten waste categories. Each annotation is associated with class-level recognition rules covering visual cues, likely confusions, and discriminative features. We evaluate six VLMs under zero-shot, two-shot, reasoning-guided, and fine-tuned settings using detection, counting, and hallucination metrics. For resource-efficient adaptation, we jointly fine-tune Qwen3-VL-2B on boxes, labels, and reasoning chains using QLoRA. Fine-tuning increases recall from 0.0248 to 0.2339 and F1 from 0.0257 to 0.2163, while reducing image-level hallucination from 0.6836 to 0.0883. However, over three-quarters of instances remain undetected, establishing WADE as a challenging benchmark for dense floating-waste grounding with compact VLMs.