arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.28216cs.CV

WALDO:杂乱场景中的一次性示例条件目标检测

WALDO: One-Shot Exemplar-Conditioned Object Detection in Cluttered Scenes

Kishor Datta Gupta, Ahmed Rafi Hasan, Md. Mahfuzur Rahman, Md. Sadman Haque, Mohd Ariful Haque

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出WALDO检测头,利用V-JEPA 2.1特征实现低成本一次性示例条件目标检测,在杂乱场景的AP@50优于Grounding DINO,验证了世界模型特征对定位和不存在检测的迁移能力。

中文摘要 AI 辅助

使用单张参考图像和简短描述在杂乱场景中定位特定目标实例,并在该实例不存在时进行报告,通常由大型视觉-语言模型来完成该任务。本文探究能否利用世界模型预训练目标已学习到的表示,以更低成本实现相同能力。我们提出WALDO,这是一个具有340万个可训练参数的一次性示例和语言条件检测头,它读取冻结的V-JEPA 2.1特征,联合预测目标定位和目标存在情况,且主干网络无需梯度。由于示例条件监督数据稀缺,我们从实例标注中合成训练 episode,从真实框中挖掘示例,并构建排除参考实例但保留同类别干扰项的不存在案例。这极易出错:在明显的实现中,仅裁剪尺寸就能预测标签,基于此训练的检测头在从未参考示例的情况下达到0.9998的不存在AUROC,我们报告了消除该捷径的负对照实验。在35个保留的杂乱场景上,WALDO达到0.461的目录AP@50,而在相同评分规则下,提示后的Grounding DINO基线仅为0.306。在匹配的576-token网格下,用DINOv3替代V-JEPA会使同类别不存在AUROC从0.880降至0.726,实例AP@50从0.201降至0.141,从而将性能提升的来源确定为预训练目标而非输入分辨率。不过,实例级Success@1仅达到0.190,而类别 chance 基准为0.190:世界模型特征可迁移到定位精度和不存在检测,但无法迁移到实例身份识别。

英文摘要

Locating a specific object instance in a cluttered scene using a single reference image and a short description, and reporting when that instance is absent, large vision-language models usually address this task. We ask whether the same capability is available far more cheaply, from representations already learned by a world-model pretraining objective. We present WALDO, a one-shot exemplar- and language-conditioned detection head with 3.4M trainable parameters that reads frozen V-JEPA 2.1 features to jointly predict object localization and target presence, with no gradient on the backbone. Because exemplar-conditioned supervision is scarce, we synthesize training episodes from instance annotations, mining exemplars from ground-truth boxes and constructing absence cases that exclude the referenced instance while leaving same-category distractors in view. This is easy to get wrong: in the obvious implementation, crop size alone predicts the label, and a head trained on it reaches 0.9998 absence AUROC without ever consulting the exemplar, and we report the negative controls that close the shortcut. On 35 held-out cluttered scenes, WALDO achieves a 0.461 catalogue AP@50, compared to 0.306 for a prompted Grounding DINO baseline under an identical scorer. Substituting DINOv3 for V-JEPA under a matched 576-token grid drops within-category absence AUROC from 0.880 to 0.726 and instance AP@50 from 0.201 to 0.141, isolating the pretraining objective rather than input resolution as the source of the gain. Instance-level Success@1, however, reaches only 0.190 against a 0.190 category-chance floor: world-model features transfer to localization precision and absence detection but not to instance identity.

发表机构

  • Clark Atlanta University(克拉克亚特兰大大学)
  • United International University(联合国际大学)

机构由 AI 辅助整理,请以论文原文为准。

↑