发表机构
UNC Chapel Hill; University of Texas at Austin; Microsoft(北卡罗来纳大学教堂山分校; 德克萨斯大学奥斯汀分校; 微软)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
TeleTune从无目标、不可重放且可能交错任务的离线遥测日志中学习文本技能库,通过动作预测误差优化技能,并利用工作流检索演示,在WorkArena和Online-Mind2Web上分别达到77.1%和80.6%的平均成功率,优于现有基线。
AI 中文摘要
计算机使用智能体需要捕获人们如何使用软件的程序性知识。用户遥测为这类知识提供了可扩展的来源。然而,从这些日志中学习可复用技能需要应对三个挑战:(1)目标欠指定,因为日志不记录每个动作背后的目标;(2)不可重放性,因为过去的活动无法重放以评估技能更新;(3)交错轨迹,因为日志可能混合多个任务而不标记其边界。为解决这些问题,我们提出了TeleTune,一个从离线日志中学习文本技能库的框架,这些日志没有记录目标,在优化过程中无法重放,并且可能交错任务。TeleTune利用日志轨迹上的动作预测误差来提出库编辑,并仅保留那些能提高留出动作预测准确性的编辑,我们称之为技能引导的进展。学习到的工作流还能检索覆盖新任务子目标的演示。在测试时,智能体获得学习到的库和基于工作流检索的演示。在WorkArena和Online-Mind2Web上的实验表明,TeleTune优于随机检索、Agent Workflow Memory(AWM)及其组合。我们发现最佳基线因设置而异,而TeleTune分别达到77.1%和80.6%的平均成功率,在每个基准上比最强基线提高6.7%和7.7%。在WorkArena训练数据最重的扰动下,TeleTune保持最高平均成功率68.5%,比最强基线高6.3%。我们的分析表明:(1)技能优化和基于工作流的检索是互补的;(2)在固定日志上优化比用实时情节验证相同编辑节省5到75倍的令牌;(3)技能引导的进展跟踪实时成功率。
英文摘要
Computer-use agents need to capture procedural knowledge of how people use software. User telemetry offers a scalable source of this knowledge. However, learning reusable skills from these logs requires addressing three challenges: (1) Goal Underspecification, since logs do not record the goal behind each action; (2) Non-Replayability, since past activity cannot be replayed to evaluate skill updates; and (3) Interleaved Trajectories, since logs may mix several tasks without marking their boundaries. To address these, we introduce TeleTune, a framework for learning a textual skill library from offline logs without recorded goals, cannot be replayed during optimization, and may interleave tasks. TeleTune uses action-prediction errors on logged trajectories to propose library edits and keep only those that improve held-out action-prediction accuracy, which we call skill-guided progress. The learned workflows also enable retrieval of demonstrations that cover the subgoals of a new task. At test time, the agent is provided with the learned library and the workflow-based retrieved demonstrations. Experiments on WorkArena and Online-Mind2Web show that TeleTune outperforms random retrieval, Agent Workflow Memory (AWM), and their combination. We find that the best baseline varies by setting, whereas TeleTune achieves average success rates of 77.1% and 80.6%, respectively, improving over the strongest baseline on each benchmark by 6.7% and 7.7%. Under the heaviest perturbation of the WorkArena training data,TeleTune keeps the highest average success rate at 68.5%, 6.3% above the strongest baseline. Our analyses show (1) skill optimization and workflow-based retrieval are complementary, (2) optimizing on fixed logs costs 5 to 75 times fewer tokens than validating the same edits with live episodes, (3) skill-guided progress tracks the live success rate.
CommentsProject Page: https://microsoft-teletune.github.io/