发表机构
Physical Superintelligence Lab, Fysics AI; College of Intelligent Robotics and Advanced Manufacturing, Fudan University; Multimedia Laboratory (MMLab), The Chinese University of Hong Kong; College of Electronic and Information Engineering, Tongji University(Fysics AI 物理超级智能实验室; 复旦大学智能机器人及先进制造学院; 香港中文大学多媒体实验室; 同济大学电子与信息工程学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出Fysiverse-3D-Vision统一框架,通过视觉-语言-几何共享表示和物体条件布局模块,从单张图像生成可执行3D场景,实现空间推理与几何重建相互增强,提升几何一致性、布局估计和物理理解。
AI 中文摘要
生成模型推动了图像条件下的3D内容创建,但从单张图像生成可控且可执行的3D场景仍然具有挑战性。现有的3D生成方法可以合成视觉上合理的物体和场景,但其空间布局估计与特定的资产生成器耦合。它们难以联合建模物体语义、度量几何和场景级空间关系,而这些对于交互式编辑、物理模拟和具身应用至关重要。我们提出Fysiverse-3D-Vision,一个统一的视觉-语言-几何框架,用于从单张图像进行生成式3D场景重建和可执行资产构建。我们建立了一个共享表示,其中空间推理和几何重建相互增强,使得物体布局的推断能够超越单个资产生成器的限制。我们的模型在统一的Transformer中整合了文本监督、语义视觉线索和几何表示,以捕获场景上下文、度量几何和物体级交互。一个物体条件布局模块在目标物体表示和全局几何特征之间执行交叉注意力,以预测物体的平移、旋转和缩放。训练过程逐步学习几何-语言对齐,引入布局推理同时保持重建能力,并通过碰撞感知优化细化物理一致性。通过将空间布局推理与资产合成分离,Fysiverse-3D-Vision为交互式场景编辑、物体级操作和可执行3D内容生成提供了适应性接口。实验表明,与现有方法相比,我们的框架在几何一致性、布局估计、渲染质量和物理属性理解方面取得了更优的性能。
英文摘要
Generative models have advanced image-conditioned 3D content creation, yet generating controllable and executable 3D scenes from a single image remains challenging. Existing 3D generative approaches can synthesize visually plausible objects and scenes, but their spatial layout estimation is coupled with specific asset generators. They struggle to jointly model object semantics, metric geometry, and scene-level spatial relationships, which are essential for interactive editing, physical simulation, and embodied applications. We propose Fysiverse-3D-Vision, a unified vision-language-geometry framework for generative 3D scene reconstruction and executable asset construction from a single image. We establish a shared representation where spatial reasoning and geometric reconstruction mutually enhance each other, allowing object layouts to be inferred beyond the constraints of individual asset generators. Our model integrates textual supervision, semantic visual cues, and geometric representations within a unified Transformer to capture scene context, metric geometry, and object-level interactions. An object-conditioned layout module performs cross-attention between target object representations and global geometric features to predict object translation, rotation, and scale. Training progressively learns geometry-language alignment, introduces layout reasoning while preserving reconstruction capability, and refines physical consistency through collision-aware optimization. By separating spatial layout reasoning from asset synthesis, Fysiverse-3D-Vision provides an adaptable interface for interactive scene editing, object-level manipulations, and executable 3D content generation. Experiments demonstrate that our framework achieves superior geometric consistency, layout estimation, rendering quality, and physical property understanding compared with existing approaches.
CommentsFysics AI Technical Report