透过人眼与机器之眼:理解视频透视扩展现实中的视角不匹配
Through Human Eyes and Machine Eyes: Understanding View Mismatch in Video See-Through Extended Reality
浏览论文内容
中文总结 AI 辅助
本研究针对VST XR系统中系统截图与用户实际可见区域不匹配的问题,通过形式化定义和Meta Quest 3测量揭示该现象,并分析其对AI集成XR系统的安全、隐私和可靠性风险。
中文摘要 AI 辅助
视频透视扩展现实(VST XR)系统通常使用头戴式设备截图或捕获的帧作为用户第一人称视觉上下文的代理。然而,系统捕获的视图与用户的有效可见视野并不一定重合:截图记录的是矩形机器可读帧,而用户的有效可见区域可能更受限且非矩形。本文研究了VST XR中这种人与系统之间的视角不匹配。我们通过定义共可见区域、仅系统可见区域和仅人类可见区域,形式化了系统捕获区域与人类可见区域之间的关系。随后,我们在Meta Quest 3上进行了试点级别的边界测量,揭示了矩形截图帧与近似人类可见边界之间的明显不匹配。基于此模型,我们分析了视角不匹配如何影响基于截图的XR感知及下游视觉-语言模型任务。通过四个代表性案例研究,我们展示了潜在风险和失败模式,包括提示注入、隐私泄露、人类不可见信息偏差以及人类可见信息缺失。我们的结果表明,视角不匹配不仅是几何伪影,还可能对集成AI的VST XR系统引入安全、隐私和可靠性方面的担忧。
英文摘要
Video see-through extended reality (VST XR) systems commonly use headset screenshots or captured frames as proxies for the user's first-person visual context. However, the system-captured view and the user's effective visible field do not necessarily coincide: a screenshot records a rectangular machine-readable frame, whereas the user's effective visible region can be more constrained and non-rectangular. This paper studies this human-system view mismatch in VST XR. We formalize the relationship between the system-captured region and the human-visible region by defining their co-visible, system-only, and human-only regions. \rev{We then conduct a pilot-level boundary measurement on Meta Quest 3, revealing a clear mismatch between the rectangular screenshot frame and the approximate human-visible boundary. Building on this model, we analyze how view mismatch can affect screenshot-based XR sensing and downstream vision-language model tasks. Through four representative case studies, we illustrate potential risks and failure modes including prompt injection, privacy leakage, human-invisible information bias, and missing human-visible information. Our results show that view mismatch is not only a geometric artifact, but can also introduce security, privacy, and reliability concerns for AI-integrated VST XR systems.
发表机构
- Duke University(杜克大学)
机构由 AI 辅助整理,请以论文原文为准。