任务自适应的基于2D VLM的接地3D程序员
Task-Adaptive Grounded 3D-Programmers Using 2D VLMs
浏览论文内容
中文总结 AI 辅助
本文提出3D-Prog框架,通过规范坐标框架(CCF)和任务自适应反馈(TAF)使2D VLM无需重训练即可执行开放词汇的3D理解、操作和生成,实现几何感知的3D编程。
中文摘要 AI 辅助
最近的视觉语言模型(VLMs)展现出显著的泛化和推理能力,然而这些模型中的3D理解受到数据规模、训练多样性和推理能力的限制。我们没有天真地将这些模型扩展到3D,而是采取了一种不同的方法:通过引入3D接地和迭代反馈循环,我们使强大的2D VLM能够在3D中可靠地操作,其中包含两个新概念:规范坐标框架(CCF)和任务自适应反馈(TAF)。CCF作为一种统一的视觉表示,将输入和输出锚定到共享的欧几里得坐标系中,解决了3D接地中的常见挑战,如轴模糊、不一致的度量尺度和浮动的参考。与这种对3D输入的结构化框架互补,TAF通过任务自适应的动态反馈闭合推理循环,使2D VLM能够在其原生视觉上下文中执行各种开放词汇任务。在此基础之上,我们引入了3D-Prog,一个3D理解、推理和生成框架,它联合使用CCF和TAF以及强大的VLM。无需任何重新训练,3D-Prog即可在对象级和场景级任务中执行开放词汇的3D理解、操作和生成。我们的实验表明,CCF和TAF的联合使用将2D VLM转变为具有几何感知的3D程序员,在多样化的3D任务中实现了一致、可解释且高质量的结果。
英文摘要
Recent vision-language models (VLMs) exhibit remarkable generalization and reasoning abilities, yet 3D understanding in these models is limited by data scale, training diversity, and reasoning capacity. Instead of naively extending these models into 3D, we take a different approach: we enable powerful 2D VLMs to operate reliably in 3D by introducing 3D grounding and iterative feedback loops with two novel concepts: Canonical Coordinate Framing (CCF) and Task-Adaptive Feedback (TAF). CCF serves as a unified visual representation that anchors both inputs and outputs to a shared Euclidean coordinate system, solving common challenges in 3D grounding such as axis ambiguity, inconsistent metric scale, and floating references. Complementary to this structured framing of the 3D inputs, TAF closes the reasoning loop with task-adaptive dynamic feedback that enables 2D VLMs to perform varied open-vocabulary tasks within their native visual context. Building on this foundation, we introduce 3D-Prog, a 3D understanding, reasoning, and generation framework that jointly employs the capabilities of CCF and TAF together with powerful VLMs. Without requiring any retraining, 3D-Prog performs open-vocabulary 3D understanding, manipulation, and generation across both object-level and scene-level tasks. Our experiments show that the joint use of CCF and TAF transforms 2D VLMs into geometry-aware 3D programmers, achieving consistent, interpretable, and high-quality results across diverse 3D tasks.
发表机构
- TextQL(TextQL公司)
- INSAIT, Sofia University “St. Kliment Ohridski”(圣克莱门特奥赫里德斯基索菲亚大学INSAIT研究所)
机构由 AI 辅助整理,请以论文原文为准。