发表机构
Saarland University; Singapore-MIT Alliance for Research and Technology; Microsoft; Microsoft Research(萨尔兰大学; 新加坡-麻省理工学院研究联盟; 微软; 微软研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出任务后工作流(可编辑的图表示)以帮助用户理解、验证和复用AI智能体的执行过程,实验表明其优于仅提示条件,且调整工作流与调整提示效果相当。
AI 中文摘要
AI智能体能够通过将单个自然语言请求转化为跨越工具、文件和应用程序的多步骤过程来自动化任务。用户往往只能从零散的执行信息和最终输出中判断该过程。为了使已完成的过程更易于理解、验证和复用,我们研究了任务后工作流:即智能体已完成执行的可编辑、基于图形的表示。我们首先分析了来自n8n的10,803个公开工作流模板,以刻画真实世界的自动化实践,随后开发了Trace2Flow,一个将智能体执行轨迹转化为交互式任务后工作流的研究探针。在一项研究中,参与者(N=20)审查了带有提示或智能体错误的智能体执行。我们发现,与仅提示条件相比,任务后工作流提高了他们的理解和错误检测能力,并且验证主要发生在用户跨多个证据来源进行交叉检查时。对于后续任务,调整工作流在成功率、时间和难度上与调整先前提示相当,并且通常更受青睐。
英文摘要
AI agents can automate tasks by turning a single natural-language request into a multi-step process spanning tools, files, and applications. Users are often left to judge that process from fragmented execution information and the final output. To make the completed process easier to understand, validate, and reuse, we investigate post-task workflows: editable, graph-based representations of an agent's completed execution. We first analyzed 10,803 public workflow templates from n8n to characterize real-world automation practice, then developed Trace2Flow, a research probe that translates agent execution traces into interactive post-task workflows. In a study, participants (N = 20) reviewed agent executions with prompt or agent errors. We found that post-task workflows improved their understanding and error detection over a prompt-only condition, and that validation succeeded mainly when users cross-checked across multiple evidence sources. For follow-up tasks, adapting the workflow matched adapting the prior prompt in success, time, and difficulty, and was often preferred.