发表机构
Seoul National University(首尔大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出CLON零样本6D位姿前端,无需新物体特定任务训练,通过语言语义记忆与提示权重优化候选框,在7个BOP-Classic-Core数据集上显著提升检测、分割及下游位姿性能。
AI 中文摘要
零样本6D位姿估计流程越来越依赖强大的下游位姿求解器,但其性能常受前端限制:物体候选框需保留部分可见的真实正例,同时拒绝语义合理的干扰项。本文提出CLON(Cue-Calibrated Linguistic Object Onboarding),一种无需对新物体进行特定任务训练的前端。给定待载入物体集的渲染模板,CLON构建用于自上而下生成候选框的语言语义记忆,以及用于校准候选框评分的物体集提示权重。语言记忆引导SAM 3生成待载入物体的高召回率候选框,提示权重在场景推理前从待载入物体集计算一次,在线评分时保持固定。在7个BOP-Classic-Core数据集上,CLON相比CNOS和SAM-6D前端,将检测AP提升8.1个百分点(pp),分割AP提升6.2个百分点,下游6D位姿AR最高提升4.1个百分点。
英文摘要
Zero-shot 6D pose estimation pipelines increasingly rely on strong downstream pose solvers, but their performance is often limited by the front-end: object proposals must preserve partially visible true positives while rejecting semantically plausible distractors. We introduce Cue-Calibrated Linguistic Object Onboarding (CLON), a front-end requiring no task-specific training for new objects. Given rendered templates of the onboarded object set, CLON constructs a linguistic semantic memory for top-down proposal generation and object-set cue weights for calibrated proposal scoring. The linguistic memory guides SAM 3 toward high-recall proposals for onboarded objects, while cue weights are computed once from the onboarded object set before scene inference and kept fixed during online scoring. On seven BOP-Classic-Core datasets, CLON improves detection AP by 8.1 percentage points (pp), segmentation AP by 6.2 pp, and downstream 6D pose AR by up to 4.1 pp over CNOS and SAM-6D front-ends.
Comments8 pages, 5 figures