CR-Refiner:用于编辑条件3D场景检索的以对象为中心的最优传输重排器
CR-Refiner: An Object-Centric Optimal Transport Reranker for Edit-Conditioned 3D Scene Retrieval
浏览论文内容
中文总结 AI 辅助
研究编辑条件3D场景检索问题,提出CR-Refiner重排器,通过冻结语言模型解析编辑、不平衡最优传输问题评分候选对象、轴条件结构先验及语言模型验证器细化结果,在不同基础检索器上提升了硬子集检索指标。
中文摘要 AI 辅助
编辑条件3D场景检索将参考3D房间与自然语言修改配对,并从语料库中检索满足该编辑的房间。先前的三项工作在这项任务上均存在不足。二维合成图像检索基于像素级编辑进行推理,对3D对象集没有原语。3D基础编码器嵌入单个对象,但无法在场景级别进行合成。3D场景定位方法在静态场景中定位参考,而不是在语料库中对修改后的房间进行排序。我们提出了CR-Refiner,一种无需训练的重排器,它用三个组件包装任何基础检索器的前K个候选对象。一个冻结的语言模型将编辑解析为结构化查询实体,每个候选对象通过在1xG成本矩阵上的不平衡最优传输问题进行评分,该矩阵耦合类别、风格、材料和几何形状。不平衡求解器让单实体查询在不相关对象上丢弃质量,直接对不对称性进行建模。轴条件结构先验为几何编辑添加尺寸关键字线索,为空间编辑添加主题锚定方向线索。一个语言模型验证器用连续置信度细化前三个候选对象。由于没有基准评估对3D对象集的合成匹配,我们还发布了3D-CER,这是一个跨越五个编辑轴的23381个房间室内语料库上的4963个编辑条件查询,具有多正真值、CIRR风格的硬子集和零目标对抗样本。在三种性质不同的基础检索器上,CR-Refiner在每个编辑轴上都持续提高了硬子集R@1和mAP@10。
英文摘要
Edit-conditioned 3D scene retrieval pairs a reference 3D room with a natural-language modification and retrieves rooms from a corpus that satisfy the edit. Three lines of prior work each fall short on this task. 2D composed image retrieval reasons over pixel-level edits and has no primitive for 3D object sets. 3D foundation encoders embed individual objects but cannot compose at the scene level. 3D scene-grounding methods localize references inside a static scene rather than rank modified rooms across a corpus. We present CR-Refiner, a training-free reranker that wraps any base retriever's top-K candidates with three components. A frozen LLM parses the edit into a structured query entity, and each candidate is scored by an unbalanced optimal-transport problem over a 1xG cost matrix coupling category, style, material, and geometry. The unbalanced solver lets the single-entity query drop mass on irrelevant objects, modelling the asymmetry directly. An axis-conditional structural prior adds size-keyword cues for geometric edits and subject-anchor direction cues for spatial edits. An LLM verifier refines the top three candidates with continuous confidence. Because no benchmark evaluates compositional matching over 3D object sets, we additionally release 3D-CER, 4,963 edit-conditioned queries over a 23,381-room indoor corpus across five edit axes, with multi-positive ground truth, CIRR-style hard subsets, and zero-target adversarials. Across three qualitatively distinct base retrievers, CR-Refiner consistently improves hard-subset R@1 and mAP@10 on every edit axis.
发表机构
- Hong Kong University of Science and Technology(香港科技大学)
- Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
机构由 AI 辅助整理,请以论文原文为准。