AI 中文总结
本研究针对沉浸式平台中多模态空间交互问题,提出基于参考的操作框架,分解空间参考为源、锚点、框架组件,实现LLM驱动的流程并验证其效用,为空间交互设计提供关键经验。
AI 中文摘要
在沉浸式平台中通过语音和手势操作物体时,用户会自然构建空间参考,涉及场景实体、自身身体或环境。本研究利用空间认知理论,系统探究用户如何构建与传达空间意图。借助定制工具包,开展了虚拟现实中场景构建的“巫师助手(Wizard-of-Oz)”研究,以观察无约束的多模态(语音+手势)输入模式。基于该发现,本研究形式化了一个框架,将空间参考分解为源(Source)、锚点(Anchor)和框架(Frame)三个核心组件,并刻画其组合策略与显式程度。通过实现基于大语言模型(LLM)的流程,结合一组示例交互技术开展初步技术评估,验证了该基于参考的操作框架的实用性。最后,讨论了支持基于参考的空间交互的关键经验教训。
英文摘要
When manipulating objects in immersive platforms through speech and gesture, users naturally construct spatial references, referring to scene entities, their bodies, or the environment. Leveraging spatial cognition theories, this work systematically examines how users construct and communicate spatial intent. Using a custom toolkit, we conducted a Wizard-of-Oz study to observe unconstrained multimodal (speech + gesture) input patterns in Virtual Reality for scene construction. Based on these findings, we formalize a framework that decomposes spatial references into three core components: Source, Anchor, and Frame, while characterizing their compositional strategies and explicitness. We demonstrate the utility of this Reference-based Manipulation framework by implementing an LLM-based pipeline featuring a set of example interaction techniques with a preliminary technical evaluation. Finally, we discuss key lessons learned for supporting reference-based spatial interaction.
Comments16 pages, 10 figures. Accepted to the 39th Annual ACM Symposium on User Interface Software and Technology (UIST 2026)