发表机构
Tianjin University; Shenzhen University of Advanced Technology; Shanghai Jiaotong University; Shanghai AI Laboratory(天津大学; 深圳理工大学; 上海交通大学; 上海人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对实时开放词汇检测中紧凑模型难以吸收丰富语义的问题,提出RT-DETR-World,通过训练时迁移LLM描述语义并保持推理轻量匹配,实现竞争性零样本精度与高效权衡。
AI 中文摘要
开放词汇检测(OVD)通过文本类别查询来识别训练中未见过的类别,然而在实现实时效率的同时保持强大的泛化能力仍然具有挑战性。除了词汇规模扩展外,零样本泛化可能受益于从已见数据中学到的可复用视觉-语义线索,包括属性、动作、状态和上下文关系。现有的实时OVD方法主要强调词汇覆盖率和高效的区域/查询-文本匹配;在严格的效率约束下,紧凑型检测器可能难以吸收丰富的实例语义和场景上下文。我们提出RT-DETR-World,一种紧凑的DETR风格检测器,它在训练期间迁移描述所传达的丰富语义,同时在推理时保持轻量级的查询-文本匹配。我们构建了具有三级监督的GroundingCapv2:用于标准OVD的类别名称、传达实例级语义的对象描述以及传达对象关系和场景上下文的图像描述。这些描述仅作为训练时的语义监督。为了帮助紧凑型检测器吸收这些语义,我们提出了双路径描述对齐(DDA),结合了部署一致的MiniLM路径和仅训练用的LLM教师。MiniLM提供查询-类别监督和对象描述对齐,而离线教师特征分别在对象和图像级别监督匹配查询和全局视觉表示。所有教师特征都是预计算的,教师侧模块在训练后被移除。我们进一步提出了关系感知负样本松弛(RNR),它利用教师派生的语义相似性来松弛相关负样本,同时保留精确正样本。实验证明了具有竞争力的零样本准确性和有利的准确性-效率权衡。代码将发布。
英文摘要
Open-vocabulary detection (OVD) recognizes categories unseen during training through textual category queries, yet achieving strong generalization with real-time efficiency remains challenging. Beyond vocabulary scaling, zero-shot generalization may benefit from reusable visual--semantic cues learned from seen data, including attributes, actions, states, and contextual relations. Existing real-time OVD methods primarily emphasize vocabulary coverage and efficient region/query--text matching; under strict efficiency constraints, compact detectors may struggle to absorb rich instance semantics and scene context. We propose RT-DETR-World, a compact DETR-style detector that transfers the rich semantics conveyed by descriptions during training while retaining lightweight query--text matching at inference. We construct GroundingCapv2 with three levels of supervision: category names for standard OVD, object descriptions conveying instance-level semantics, and image descriptions conveying object relations and scene context. These descriptions serve only as training-time semantic supervision. To help the compact detector absorb these semantics, we propose Dual-Path Description Alignment (DDA), combining a deployment-consistent MiniLM pathway with a training-only LLM teacher. MiniLM provides query--category supervision and object-description alignment, while offline teacher features supervise matched queries and global visual representations at the object and image levels, respectively. All teacher features are precomputed, and the teacher-side modules are removed after training. We further propose Relation-Aware Negative Relaxation (RNR), which uses teacher-derived semantic similarities to relax related negatives while preserving exact positives. Experiments demonstrate competitive zero-shot accuracy and a favorable accuracy--efficiency trade-off. The code will be released.