arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.05777cs.AIcs.CL

从生产轨迹中挖掘智能体技能

Mining Agent Skills from Production Traces

  • Massachusetts Institute of Technology(麻省理工学院)
  • Microsoft Corporation(微软公司)

机构由 AI 辅助整理,请以论文原文为准。

Yue Ran Kang, Colton Mikolajczyk, Chhaya Methani, Hazel Mak, Sahil Bhatnagar, Susheel Suresh, Alejandro Gutierrez Munoz

AI总结:

本研究比较不同挖掘证据和技能形式对智能体技能性能的影响,发现其效果依赖企业领域,建议按任务需求定制元技能而非一刀切。

AI中文摘要:

记录程序性指令的智能体技能越来越多地从执行轨迹中挖掘,而非手工整理。技能挖掘流程通常利用已知的任务结果或反馈来指导技能构建。在生产环境中,关于一次运行是否成功的可靠信息可能无法获得。我们研究了执行轨迹的采样、成功或失败信息的获取以及挖掘技能的形式如何影响下游任务性能。在保持挖掘流程不变的情况下,我们比较了挖掘证据和技能形式的六种组合。挖掘证据有三个层次:仅成功轨迹;成功和失败及其结果标签;或相同混合但标签被隐藏。技能形式有两种类型:有序的工作流计划,或实体、状态和策略的声明性本体。我们在两个企业基准测试ThinkingBox-Bench和APEX-Agents上评估了挖掘的技能。对任务级配对差异的分析表明,不同挖掘证据和技能形式配置的收益取决于企业领域。在ThinkingBox-Bench上,配对差异显示工作流比本体得分高1.7个百分点,Goldilocks比仅成功证据类型高2.4个百分点,而Goldilocks盲模拟(在没有结果的情况下学习的技能)差3.1个百分点。APEX-Agents显示出对本体的适度偏好,并且在证据制度之间没有明确偏好。在每个领域内,与任务结构相关的约束导致挖掘技能的性能不均匀。这些发现激励我们根据目标任务的需求定制元技能,而不是采用一刀切的方法。

英文摘要:

Agent skills that record procedural instructions are increasingly mined from execution traces rather than curated by hand. Skill-mining pipelines often use known task outcomes or feedback to guide skill construction. In production, reliable information on whether a run has succeeded may be unavailable. We study how the sampling of execution traces, access to success or failure information, and the form of the mined skills affect downstream task performance. Holding the mining pipeline fixed, we compare six combinations of mining evidence and skill forms. Mining evidence has three levels: successful trajectories only, successes and failures with their outcome labels, or the same mix with labels withheld. Skill form has two types: an ordered workflow plan, or a declarative ontology of entities, states, and policies. We evaluate the mined skills on two enterprise benchmarks, ThinkingBox-Bench and APEX-Agents. Analysis of task-level paired differences shows that the benefits of different configurations of mining evidence and skill forms depend on the enterprise domain. On ThinkingBox-Bench, paired differences show that workflows score better than ontology by 1.7 pp, Goldilocks beats success-only evidence type by 2.4 pp and Goldilocks blind simulating skills learnt without outcomes is worse by 3.1 pp. APEX-Agents shows a moderate preference for ontologies and no clear preference between evidence regimes. Within each domain, task structure related constraints drive uneven performance with mined skills. These findings motivate tailoring meta-skills to the demands of the target tasks rather than adopting a one-size-fits-all approach.

补充信息

↑