arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.20621cs.CV

RECOUNT:基于合成视觉示例的参考引导计数

RECOUNT: Reference-guided Counting with Synthetic Visual Exemplars

Adriano D'Alessandro, Ali Mahdavi-Amiri, Ghassan Hamarneh

首次发表
浏览论文内容

中文总结 AI 辅助

RECOUNT是基于合成视觉示例的图像引导零样本计数即插即用框架,通过扩散模型扩展参考图像生成示例库,在两个基准上实现最优零样本计数准确率,大幅降低计数误差。

中文摘要 AI 辅助

文本引导的零样本对象计数模型擅长空间定位,但对新颖或细粒度类别的分类效果较差:自然语言过于粗糙,无法完全指定视觉身份,因此无法区分视觉相似的干扰项。少样本计数模型通过视觉示例规避了这一问题,但需要对每张图像进行手动标注。为解决这一难题,我们提出RECOUNT,一个用于图像引导零样本计数的即插即用框架。与用文本提示指定类别不同,我们的核心见解是从一张离线参考图像中以视觉方式指定类别。然而,我们发现单一参考图像对类别的外观覆盖范围较窄,在不同场景下不可靠。因此,我们将扩散模型重新用作自动对比数据引擎,将参考图像扩展为多样化的示例库,提供文本无法提供的判别性细节。RECOUNT保留了任何冻结计数模型的类不可知提议,并将分类任务卸载到单独的视觉模块(一个冻结的骨干网络,带有在该合成数据上训练的轻量级头部),该模块将每个提议与目标和干扰项库进行匹配。应用于冻结计数模型时,RECOUNT在两个基准上达到了最佳零样本准确率,与最强的现有零样本计数模型相比,在LookAlikes上的计数误差(MAE)降低了55%,在PairTally上降低了21%。

英文摘要

Text-guided zero-shot object counters excel at spatial localization but categorize poorly on novel or fine-grained classes: natural language is too coarse to fully specify visual identity, so they fail to separate visually similar distractors. Few-shot counters sidestep this with visual exemplars, but require manual annotations on every image. To resolve this dilemma, we introduce RECOUNT, a plug-and-play framework for image-guided zero-shot counting. Rather than specify a category with a text prompt, our key insight is to specify it visually, from a single off-scene reference image. However, we find that a lone reference image provides narrow coverage of a category's appearance and is unreliable across diverse scenes. We therefore repurpose a diffusion model as an automated contrastive data engine that expands the reference into a diverse exemplar gallery, supplying the discriminative detail that text cannot. RECOUNT preserves the class-agnostic proposals of any frozen counter and offloads categorization to a separate visual module (a frozen backbone with a lightweight head trained on this synthetic data) that matches each proposal against the target and distractor galleries. Applied to a frozen counter, RECOUNT attains the best zero-shot accuracy on both benchmarks, cutting counting error (MAE) by 55% on LookAlikes and 21% on PairTally relative to the strongest prior zero-shot counter.

发表机构

  • Simon Fraser University(西蒙菲莎大学)

机构由 AI 辅助整理,请以论文原文为准。

↑