arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.30120cs.SE

评估智能体技能在特定版本插件迁移中的表现:一项回顾性研究

Evaluating Agent Skills for Version-Specific Plugin Migration: A Retrospective Study

Beiming Liu, Haihao Li, Minjie Chen, Ning Chen, Yiran Wang, Jiming Ye, Puzhao Zhang, Tongtao Wang, Sheng Gao, William Jin, Weihao Mu, Chengzhi Liu, Yucheng Xia,… 展开作者

Beiming Liu, Haihao Li, Minjie Chen, Ning Chen, Yiran Wang, Jiming Ye, Puzhao Zhang, Tongtao Wang, Sheng Gao, William Jin, Weihao Mu, Chengzhi Liu, Yucheng Xia, Guangren Wang, Chaoyang Fan, Changfeng Huang, Xunming Lin, Yuanjie Shen

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过回顾性分析评估插件迁移技能,发现诊断分数提升但存在评分错误,提出可追溯评估方法连接聚合奖励与契约级证据,并给出审查检查建议。

中文摘要 AI 辅助

智能体技能为编码智能体打包了特定版本的维护知识,但更高的诊断分数本身并不能证明由此产生的迁移建议满足目标版本的契约。我们通过一个包含16个静态迁移任务、每个条件下两次尝试、共328个标准决策的64份报告档案,研究了一项已发布的插件升级技能。使用该技能后,平均记录奖励从93.83升至98.75,增益为4.92分(95%任务自助法区间[0.31, 10.86]);该增益集中于一个任务,且有八个任务对达到上限。将每个决策追溯至其契约领域并深入审查十份报告,揭示了偏向任一方的评分错误;在其中一份报告中,一个接受父目录的包含谓词仍获得满分。可执行探针证实了这一缺陷,并表明一个有效的拆卸修复仅因更窄的生命周期规则而被排除。替换所审查的决策后,估计值保持为正(4.61至5.39分),但其区间移至零或跨越零。使用来自另外两个模型家族的评判员,在不提供臂标签或先前分数的情况下,对所有64份报告重新评分,与原评判员在91.8%和95.7%的决策上达成一致(加权κ=0.64和0.72),并给出10.63和6.09分的增益。该研究提供了一种可追溯的评估方法,将聚合奖励与契约级证据和评判员敏感性联系起来,并为迁移建议提供了具体的审查检查。可执行的端到端修复、独立的人工标注以及其他框架留待未来工作。

英文摘要

Agent skills package version-specific maintenance knowledge for coding agents, but a higher diagnostic score does not by itself show that the resulting migration advice satisfies the target version's contract. We study a shipped plugin-upgrade skill through an archive of 64 reports on 16 static migration tasks, with two attempts per condition and 328 criterion decisions. With the skill, mean recorded reward rises from 93.83 to 98.75, a gain of 4.92 points (95% task-bootstrap interval [0.31, 10.86]); the gain is concentrated in one task, and eight task pairs are at the ceiling. Tracing every decision to its contract domain and reviewing ten reports in depth exposes grading errors that favor either arm; in one, a containment predicate that accepts the parent directory still receives full credit. Executable probes confirm this defect and show that a working teardown repair is excluded only by a narrower lifecycle rubric. Replacing the reviewed decisions keeps the estimate positive (4.61 to 5.39 points) but moves its interval to or across zero. Re-grading all 64 reports with judges from two other model families, without arm labels or prior scores, agrees with the original judge on 91.8% and 95.7% of decisions (weighted $κ=0.64$ and $0.72$) and gives gains of 10.63 and 6.09 points. The study contributes a traceable evaluation that connects aggregate reward to contract-level evidence and judge sensitivity, together with concrete review checks for migration advice. Executable end-to-end repairs, independent human annotation, and other frameworks are left to future work.

发表机构

  • Tsinghua University(清华大学)
  • Fudan University(复旦大学)
  • PetroChina Southwest Oil & Gasfield Company(中国石油西南油气田公司)
  • Jilin University(吉林大学)
  • The Frederick Gunn School(弗雷德里克·冈学校)
  • Shenzhen University(深圳大学)
  • Dalian Neusoft University of Information(大连东软信息学院)
  • University of Chinese Academy of Sciences(中国科学院大学)
  • Dalian University of Technology(大连理工大学)
  • Jiyin Zhiyuan (Shanghai) Technology Co., Ltd.(基音致远(上海)科技有限公司)
  • Lanzhou University of Technology(兰州理工大学)
  • Harbin Engineering University(哈尔滨工程大学)
  • Alibaba Cloud(阿里云)
  • Jianghan University(江汉大学)
  • Jiuxiangxian (Beijing) Technology Co., Ltd.(九象先(北京)科技有限公司)
  • Sun Yat-sen University(中山大学)
  • Great Bay University(东莞理工学院)
  • Beijing Information Science and Technology University(北京信息科技大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑