LegacyWorld:面向遗留工作流的GUI智能体的原子性感知评估
LegacyWorld: Atomicity-Aware Evaluation of GUI Agents for Legacy Workflows
AI总结:
该研究开发LegacyUse框架以自动化遗留工作流,构建含28个Windows GUI工作流的基准,采用原子性评估6款计算机使用智能体,发现不同操作特征并提出相关首要要求。
AI中文摘要:
遗留及类遗留企业系统常因关键工作流暴露的可编程接口有限、仍需手动GUI交互而难以现代化。本文报告一项部署前评估研究,其动机是开发LegacyUse——一款面向行业的、用于通过多模态LLM智能体自动化此类工作流的框架。在框架开发期间,领域专家协助识别了有状态工作流:成功的演示并不足够,智能体运行失败仍可能在业务或医疗记录中留下持续的无效变更。因此,我们采用原子性评估计算机使用智能体:一次运行要么正确完成预期工作流,要么失败且无意外的持续副作用。我们构建了由领域专家参与的基准,包含28个Windows GUI工作流,每个工作流均指定了初始状态、目标状态和特定任务的验证器。我们将专家精心设计的提示与从专家黄金路径执行的屏幕录制生成的提示进行比较。在六个托管计算机使用智能体上,我们的结果显示,有用完成、安全失败和非原子副作用是不同的操作特征。我们得出结论,工作流捕获、状态验证器和原子性感知验收测试应成为基于AI的遗留工作流自动化的首要要求。
英文摘要:
Legacy and legacy-like enterprise systems often remain difficult to modernize because critical workflows expose limited programmable interfaces and still require manual GUI interaction. This paper reports a pre-deployment evaluation study motivated by the development of legacy-use, an industry-oriented framework for automating such workflows with multimodal LLM agents. During framework development, domain experts helped identify stateful workflows where successful demos are not sufficient: a failed agent run may still leave persistent invalid changes in business or healthcare records. We therefore evaluate computer-use agents using atomicity: a run should either complete the intended workflow correctly or fail without unintended persistent side effects. We construct a domain-expert-informed benchmark of 28 Windows GUI workflows, each specified with an initial state, goal state, and task-specific validator. We compare expert-crafted prompts with prompts generated from screen recordings of expert golden-path executions. Across six hosted computer-use agents, our results show that useful completion, safe failure, and non-atomic side effects are distinct operational profiles. We conclude that workflow capture, state validators, and atomicity-aware acceptance tests should be first-class requirements for AI-based legacy workflow automation.