arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

生成式教程:面向物理任务的实时情境化视觉指令

Generative Tutorial: Towards Live Contextualized Visual Instructions for Physical Tasks

Muzhe Wu, Zuchen Li, Xu Wang, Anhong Guo

arXiv 2609.24955首次发表:更新:

发表机构

University of Michigan(密歇根大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出生成式教程框架,利用增强现实和生成模型,根据用户环境实时生成视觉指令,实验表明能提升任务执行质量并缩短确认时间。

AI 中文摘要

物理任务的视觉指令通常在一个情境中编写,而在另一个情境中执行,这要求用户将演示的工具、材料和空间关系转换到自己的环境中。我们提出了生成式教程(Generative Tutorial),一个用于实时视觉指令的概念框架,该框架在用户的环境和任务流程中描绘预期结果和动作。对最先进的图像和视频生成模型进行的一项形成性评估,识别了在15个物理任务中的失败和潜在益处。基于这些发现,我们构建了一个增强现实原型系统,该系统利用观察到的工作区上下文和先前动作的预测视觉结果,主动生成目标图像和演示视频。一项有24名参与者参与的实验室研究发现,与预先编写的指导相比,使用该系统时任务执行质量更高、感知的工作区对应性更强、步骤确认间隔更短。定性研究结果强调了情境相似性如何影响信任、生成错误如何影响理解,以及指导传递应如何适应用户需求,为未来设计提供了信息。

英文摘要

Visual instructions for physical tasks are typically authored in one context and followed in another, requiring users to translate demonstrated tools, materials, and spatial relationships into their own environment. We introduce Generative Tutorial, a conceptual framework for live visual instruction that depicts intended outcomes and actions within the user's environment and task flow. A formative evaluation of state-of-the-art image and video generation identifies failures and potential benefits across 15 physical tasks. Drawing on these findings, we build an augmented-reality prototype system that proactively generates goal images and demonstration videos using observed workspace context and predicted visual outcomes of preceding actions. A 24-participant lab study found higher task performance quality, greater perceived workspace correspondence, and shorter step-confirmation intervals with the system than with pre-authored guidance. Qualitative findings highlighted how contextual resemblance shapes trust, how generation errors affect interpretation, and how guidance delivery should adapt to users' needs, informing future designs.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑