arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FuncRoom-Agent:序贯前馈式3D功能室内场景生成

FuncRoom-Agent: Sequential Feed-Forward 3D Functional Indoor Scene Generation

Hao Feng, Zhi Zuo, MingJian Liang, Jingyu Hu, Xiaowei Hu, Liupengfei Wu, Dian Zhang, Guoxin Fang, Zhengzhe Liu

arXiv 2608.29519首次发表:更新:

AI 中文总结

FuncRoom-Agent提出序贯前馈框架,结合递归DSL与ScenePRM奖励,实现高效高质量3D功能室内场景生成,在两类生成任务上达SOTA性能。

AI 中文摘要

我们提出了功能房间生成这一全新的室内3D场景生成设置,旨在创建具备明确功能目标的房间,而非仅生成视觉上合理的布局。现有的智能体与可执行方法虽提升了可控性,但往往依赖成本高昂的测试时生成-评估-修正循环,导致功能房间生成过程缓慢且计算成本高。我们通过三项技术贡献解决这一挑战:其一,设计了一种递归领域特定语言,用于有效组织功能房间所需的分层对象组合,涵盖房间结构、主要家具、密集支撑面及嵌套小对象,将房间表示为具备明确几何与功能关系的分阶段可执行程序;其二,提出序贯前馈场景构建框架,将递归构建轨迹提炼为场景构建专家,推理阶段该专家逐阶段编写可执行DSL代码,确定性执行器直接实例化各阶段,无需教师智能体、在线评论者或迭代修正;其三,引入ScenePRM这一基于执行的过程奖励框架,通过融合功能、几何、关系及未来可构建性反馈的强化学习来优化该专家。我们还建立了面向功能的基准,在通用室内场景生成与功能房间生成任务上均实现了SOTA性能,在功能完整性、关系正确性、几何可执行性及生成效率上表现更优。

英文摘要

We introduce Function-Room Generation, a new indoor 3D scene generation setting that creates rooms supporting explicit functional goals rather than merely visually plausible layouts. Existing agentic and executable methods improve controllability, but often depend on costly test-time generate--evaluate--revise loops, making functional room generation slow and computationally expensive. We address this challenge with three technical contributions. First, we design a recursive domain-specific language to effectively organize the hierarchical object compositions required by functional rooms, from room structure and major furniture to dense support-surface and nested small objects. It represents rooms as staged executable programs with explicit geometric and functional relations. Second, we propose a sequential feed-forward scene construction framework that distills recursive construction traces into a scene construction expert. At inference time, the expert writes executable DSL code stage by stage, and a deterministic executor directly instantiates each stage without teacher agents, online critics, or iterative repair. Third, we introduce ScenePRM, an execution-grounded process reward framework that improves the expert through reinforcement learning with functional, geometric, relational, and future-constructability feedback. We further establish a function-oriented benchmark and show state-of-the-art performance on both general indoor scene generation and function-room generation, achieving stronger functional completeness, relation correctness, geometric executability, and generation efficiency.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑