arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.13539cs.CVcs.GR

ThinkBLOX:基于渐进推理的3D室内场景生成

ThinkBLOX: 3D Indoor Scene Generation with Progressive Reasoning

Yuan Xiao, Can Wang, Xiangyu Kong, Jing Liao

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对传统与现有3D室内场景生成方法不足,提出基于VLM的渐进推理框架ThinkBLOX,构建数据集并采用监督微调与新强化学习方案,在多方面优于基线,支持多样3D场景应用。

中文摘要 AI 辅助

传统图形方法常以自回归或分层方式合成3D室内场景,而基于视觉语言模型(VLM)的生成器多采用一次性范式,交互式编辑时需全局重新优化或完全重建,易导致布局不佳。本文提出ThinkBLOX,一种基于VLM的渐进推理框架,将布局生成视为状态条件下的逐步推理和行动过程。构建了ThinkBLOX-Data-200K数据集,通过监督微调让VLM学习弥合推理-行动差距。还引入Tier-Decoupled GDPO解决奖励冲突问题。实验表明ThinkBLOX在物理合理性、语义对齐和交互式可编辑性方面显著优于基线,支持多种3D场景应用。

英文摘要

While traditional graphics methods often synthesize 3D indoor scenes autoregressively or hierarchically, recent vision-language model (VLM)-based generators predominantly adopt a one-shot paradigm where the full layout is planned at once. This one-shot approach often requires global re-optimization or complete reconstruction during interactive editing (e.g., inserting or moving objects) and can lead to physically or semantically poorly organized arrangements. To address these challenges, we propose ThinkBLOX, a VLM-based progressive reasoning framework that iteratively designs and refines 3D scenes. ThinkBLOX treats layout generation as a state-conditioned, step-by-step reasoningand-action process. To power this, we construct the ThinkBLOX-Data-200K dataset, containing 224,757 procedural placement pairs annotated with multi-view scene context, explicit Chain-of-Thought (CoT) rationales, and structured JSON layouts. Through supervised fine-tuning (SFT) on this dataset, the VLM learns to bridge the reasoning-action gap under incremental updates. Furthermore, recognizing that scene synthesis is inherently a multisolution task where SFT suffers from reward conflict, we introduce Tier-Decoupled GDPO. This reinforcement learning scheme organizes heterogeneous rewards into distinct tiers, stabilizing policy optimization across physical validity, semantic plausibility, and reasoning-action consistency. Extensive experiments show that ThinkBLOX significantly outperforms recent one-shot and iterative baselines in physical plausibility, semantic alignment, and interactive editability. Additionally, we show that it supports diverse applications, including both global and local generation and rearrangement of 3D scenes.

发表机构

  • City University of Hong Kong(香港城市大学)
  • University of Hong Kong(香港大学)
  • Beijing Information Science and Technology University(北京信息科技大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑