arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.16011cs.CLcs.CV

ReRef-3D:空间指代表达引导的3D场景重排基准

ReRef-3D: A Benchmark for Spatial Referring Expression-Guided 3D Scene Rearrangement

Mary Lynn Martin, Yifei Zhang, Martha Palmer, Maria Leonor Pacheco

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出ReRef-3D基准,针对3D场景语言引导放置任务,经微调的LLaVA-3D等模型在该基准上验证了关系满足度优于物理有效性等结论。

中文摘要 AI 辅助

我们提出ReRef-3D,这是一个用于3D场景中语言引导放置的基准。它包含来自998个CLEVR衍生场景的33826条指令,涵盖16种放置类型及直接、单跳、双跳引用。每条指令需被解析为有效的新放置位置。鉴于指令定义的是可接受放置的区域而非单个坐标,我们的评估会将预测结果插入场景,重新计算关系并测试关系满足度与物理有效性。每条指令还包含经验证的自然语言改写。微调后,LLaVA-3D、3D-LLM和PlaceIt3D分别能为68.3%、31.6%、22.4%的指令生成有效放置。所有模型中,关系满足度优于物理有效性;最近、之间等关系最难;表述对性能影响极小。

英文摘要

We introduce ReRef-3D, a benchmark for language-guided placement in 3D scenes. It contains 33,826 instructions across 998 CLEVR-derived scenes, spanning 16 placement families and direct, one-hop, and two-hop references. Each instruction must be resolved into a valid new placement position. Given that an instruction defines a region of acceptable placements rather than one coordinate, our evaluation inserts a prediction into the scene, recomputes relations, and tests relation satisfaction and physical validity. Each instruction also includes a verified naturalized rewrite. After fine-tuning, LLaVA-3D, 3D-LLM, and PlaceIt3D produce valid placements for 68.3%, 31.6%, and 22.4% of instructions, respectively. Across models, relation satisfaction surpasses physical validity, relations such as nearest and between are the most difficult, and phrasing has minimal effect on performance.

发表机构

  • University of Colorado Boulder(科罗拉多大学博尔德分校)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑