AI 中文总结
研究针对计算机使用代理文本反馈问题,开发Sidekick原型系统,在不同交互阶段用多模态反馈传达其状态,经30人参与的研究验证其能显著提升多任务性能,有效支持进度感知等,还展示了应用前景及对人机协作的意义。
AI 中文摘要
计算机使用代理(CUAs)可在图形用户界面内自主执行复杂的多步骤任务,通过并行多任务提高效率。然而,对CUA专家和生成式人工智能用户的初步研究表明,当前反馈主要基于文本,需要持续关注以监控进度,且追溯过去图形用户界面交互的可见性有限。基于这些发现,我们开发了一个原型系统Sidekick,用于在不同交互阶段通过多模态反馈传达CUAs的状态:当CUAs在后台运行时,Sidekick通过环境线索发出其执行状态信号;恢复与CUAs交互时,Sidekick提供已完成操作的多模态总结以支持快速恢复上下文;当CUAs在前台运行时,Sidekick通过将代理的推理语言化和可视化来提高透明度。一项有30名参与者的研究表明,与以典型聊天或环境显示形式呈现文本反馈的基线系统相比,Sidekick显著提高了与CUAs的多任务性能。Sidekick更有效地支持了进度感知、错误和操作可追溯性。最后,我们通过几个示例应用展示了Sidekick的前景,并讨论了对长期人机协作的影响。
英文摘要
Computer Use Agents (CUAs) can autonomously execute complex, multi-step tasks within GUIs, enhancing efficiency through parallel multitasking. However, our formative studies with CUA experts and GenAI users indicated that current feedback is primarily text-based, requiring sustained attention to monitor progress and offering limited visibility to trace past GUI interactions. Based on the findings, we developed a prototype system, Sidekick, for communicating CUAs' status with multimodal feedback across different stages of interaction: (i) When CUAs run in the background, Sidekick signals its execution state through ambient cues. (ii) Upon resuming interaction with CUAs, Sidekick provides multimodal summaries of completed actions to support rapid context resumption. (iii) When CUAs operate in the foreground, Sidekick enhances transparency by verbalizing and visualizing the agent's reasoning. A study with 30 participants demonstrated that Sidekick significantly improved multitasking performance with CUAs compared to baseline systems that presented textual feedback either in a typical chat or in an ambient display. Sidekick supported progress awareness, and error and action traceability more effectively. Finally, we demonstrate the promise of Sidekick through several example applications, and discuss implications for long-horizon human-agent collaboration.