arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

OPUS:一个简单而有效的开放词汇检测统一框架

OPUS: A Simple yet Effective Unified Framework for Open-Vocabulary Detection

Xiaoyan Wei, Zhimin Yao, Ruilin Yang, Wei Zhang, Yong Dai, Yi Zhang, Wei Ge

arXiv 2608.30247首次发表:更新:

发表机构

Megvii Technology Inc.; X-Humanoid(旷视科技; X-Humanoid)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出 OPUS 框架,以简单设计结合语义视觉表示与 ICA 训练策略,在 COCO 等数据集上实现最优 Visual-I 性能,平衡多类型提示精度,将混合提示转为互补。

AI 中文摘要

近期的统一开放词汇检测(OVD)支持异构提示,包括文本查询、视觉示例及其组合,但往往依赖日益复杂的设计,如繁重的跨模态融合、分阶段训练和迭代注释流程。本文重新探讨在基础模型更强大的时代,这种复杂性是否必要。研究发现,通过语义丰富的视觉表示和可扩展的 grounding 监督,统一 OVD 可以大幅简化。本文提出 OPUS(Open-vocabulary, Prompt-Unified, Simple),这是一个统一检测器,在单个框架内支持文本、交互式视觉、通用视觉及混合提示。OPUS 采用简单的三部分设计:模型架构结合基于 DINOv3-ConvNeXt-B 骨干网的语义丰富视觉编码器,该编码器具备高效混合编码,以及提示感知解码器,避免了针对特定提示的分支以实现统一提示推理。OPUS 采用单阶段文本-视觉训练策略,结合实例级对比对齐(ICA),并由基于 SAM3 的单遍数据引擎提供异构 grounding 监督。在 COCO、LVIS-minival 和 ODinW35 上的实验表明,OPUS 达到了最先进的 Visual-I 性能,分别取得 68.1/69.2/54.7 AP,同时保持文本和 Visual-G 精度的平衡。OPUS 还将混合提示从干扰转化为互补,比单独使用文本或视觉提示表现更优。这些结果表明,简单性与强大的统一提示能力可以同时实现。

英文摘要

Recent unified open-vocabulary detection (OVD) supports heterogeneous prompts, including text queries, visual exemplars, and their combinations, but often rely on increasingly complex designs such as heavy cross-modal fusion, staged training, and iterative annotation pipelines. We revisit whether such complexity is necessary in the era of stronger foundation models. Our finding is that unified OVD can be made substantially simpler with semantic-rich visual representations and scalable grounding supervision. We present OPUS (\textbf{O}pen-vocabulary, \textbf{P}rompt-\textbf{U}nified, \textbf{S}imple), a unified detector supporting text, interactive visual, generic visual, and mixed prompting within one framework. OPUS adopts a simple three-part design. Its model architecture combines a semantic-rich visual encoder, built on a DINOv3-ConvNeXt-B backbone with efficient hybrid encoding, with a prompt-aware decoder that avoids prompt-specific branches for unified prompt reasoning. OPUS is trained with a one-stage text-visual training strategy with Instance-level Contrastive Alignment (ICA), and is supported by a SAM3-based single-pass data engine for heterogeneous grounding supervision. Experiments on COCO, LVIS-minival, and ODinW35 show that OPUS achieves state-of-the-art Visual-I performance, reaching 68.1/69.2/54.7 AP, while maintaining balanced Text and Visual-G accuracy. OPUS also turns mixed prompting from interference into complementarity, improving over text or visual prompt alone. These results show that simplicity and strong unified prompting capability can be achieved together.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑