arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AutoCue:多模态大语言模型辅助将隐式输入外部化为录屏教程中的教学视觉提示

AutoCue: Multimodal LLM-Assisted Externalization of Implicit Inputs as Instructional Visual Cues in Screencast Tutorials

Shengyang Luo, Shengyao Luo, Xiaolei Guo, Fengze Zhang, James Liang, Yingjie Victor Chen

arXiv 2608.04910首次发表:更新:

AI 中文总结

AutoCue是多模态大语言模型辅助的教程增强流水线,可将录屏教程中隐式输入外部化为视觉提示,在Autodesk Maya上的评估显示其能缩短任务时间、减少交互中断并提升学习者体验

AI 中文摘要

教程视频被广泛用于学习功能丰富的软件,但实际跟随录屏教程时经常会出现问题。通过调查和情境探究,我们发现学习者经常倒回或卡住,因为教程中缺少输入元数据,关键输入信息(尤其是鼠标操作和带修饰键的操作)往往是隐式的或缺失的。为解决该问题,我们提出AutoCue,一种多模态大语言模型辅助、人在回路的教程增强流水线,用于将隐式输入外部化为教学视觉提示。AutoCue整合帧间视觉变化、旁白信号以及官方软件手册中的操作指南,以推断可能的鼠标和按键修饰操作,随后生成对齐的提示层和可编辑工件供人优化。基于多媒体学习和认知负荷理论,我们进一步开发了一种视觉提示语法,用于表示软件学习教程中的鼠标、键盘及组合输入。我们在Autodesk Maya中实例化并评估AutoCue,将自动推理聚焦于选定的经UI介导且具有可观察视觉或文本反馈的交互,同时通过可编辑创作工件支持更模糊的状态变化。在包含24名参与者的被试间研究中,经AutoCue增强的教程缩短了任务完成时间和交互中断次数,并提升了学习者报告的体验。

英文摘要

Tutorial videos are widely used for learning feature-rich software, yet following screencast tutorials often breaks down in practice. Through a survey and contextual inquiry, we found that learners frequently rewind or get stuck because critical input information, especially mouse actions and keyboard-modified operations, is often implicit or missing in tutorials without input metadata. To address this problem, we present AutoCue, a multimodal LLM-assisted, human-in-the-loop tutorial augmentation pipeline for externalizing implicit inputs as instructional visual cues. AutoCue integrates frame-to-frame visual changes, narration signals, and operation guidance from official software manuals to infer likely mouse and key-modifier actions, then produces aligned cue layers and editable artifacts for human refinement. Grounded in multimedia learning and cognitive load theory, we further develop a visual cue grammar for representing mouse, keyboard, and combined inputs in software-learning tutorials. We instantiate and evaluate AutoCue in Autodesk Maya, focusing automatic inference on selected UI-mediated interactions with observable visual or textual feedback while supporting more ambiguous state changes through editable authoring artifacts. In a between-subjects study with 24 participants, the AutoCue-augmented tutorial reduced task completion time and interaction breakdowns and showed improved learner-reported experience.

CommentsAccepted to Graphics Interface 2026 (GI 2026)

DOI:10.1145/3831423.3831452

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑