发表机构
Sungkyunkwan University; The Hong Kong University of Science and Technology(成均馆大学; 香港科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
SPHERE通过LLM增强的空间偏好学习和人机协同强化学习,将VR室内场景生成从孤立合成转变为持续协作,显著减少纠正编辑和身体需求,生成几何鲁棒且符合用户画像的布局。
AI 中文摘要
尽管大语言模型(LLMs)推动了3D室内场景合成的发展,但现有流程无法跨会话保留用户特定偏好,使得沉浸式创作成为重复且身体疲劳的过程。我们提出SPHERE,一个自适应VR生成框架,将孤立合成转变为持续的人机协同创作。SPHERE从自然多模态交互(语音和控制器编辑)中提取持久空间偏好。为确保对空间扭曲的几何鲁棒性,它将原始编辑抽象为分层约束,同时建模局部功能和全局拓扑上下文。此外,一种人机协同强化学习机制根据用户最终编辑的场景动态更新检索策略。一项混合设计用户研究(N=42)和离线消融实验表明,SPHERE显著减少了纠正性编辑和身体需求,防止偏向浅层对象级特征,从而生成几何鲁棒、与用户画像对齐的布局。最终,SPHERE展示了如何通过捕获演示的空间逻辑实现受控空间适应,为沉浸式创作建立了可靠、受治理的人机协作框架。项目页面和源代码将在此https URL提供。
英文摘要
While Large Language Models (LLMs) advance 3D indoor scene synthesis, current pipelines fail to retain user-specific preferences across sessions, making immersive authoring a repetitive and physically fatiguing process. We present SPHERE, an adaptive VR generation framework that transforms isolated synthesis into continuous human-AI co-creation. SPHERE extracts persistent spatial preferences from natural multimodal interactions (speech and controller edits). To ensure geometric resilience against spatial distortions, it abstracts these raw edits into hierarchical constraints modeling both local functional and global topological contexts. Furthermore, a human-in-the-loop reinforcement learning mechanism dynamically updates retrieval policies based on the user's final edited scenes. A mixed-design user study ($N=42$) and an offline ablation demonstrate that SPHERE significantly reduces corrective edits and physical demand, preventing bias toward shallow object-level traits to yield geometrically resilient, profile-aligned layouts. Ultimately, SPHERE demonstrates how capturing demonstrated spatial logic enables controlled spatial adaptation, establishing a reliable, governed human-AI collaboration framework for immersive authoring. Project page and source code will be available at: https://github.com/hyeonmin11/SPHERE