arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

自监督技能优化

Self-Supervised Skill Optimization

Siran Peng, Cuiyu Yang, Tianyu Fu, Tianshuo Zhang, Haoyuan Zhang, Weisong Zhao, Anyang Su, Minghui Wu, Huiying Li, Xiangyu Zhu, Chenxu Zhao, Zhen Lei

arXiv 2607.28777首次发表:更新:

AI 中文总结

该研究提出自监督技能优化(SSO)框架,仅用未标注任务实例学习可复用技能,在封闭、开放任务上优于无真实标注的提示优化器,部分场景表现接近最强的基于真实标注的技能优化器。

AI 中文摘要

智能体技能为冻结的大语言模型(LLM)智能体提供可复用的过程性指导,近期研究表明这类技能可通过真实标注(GT)反馈进行优化。然而,许多应用场景缺乏真实标注、任务分数、奖励或可靠的任务专用评估器。因此,我们引入自监督技能优化(SSO),这是一种仅从未标注任务实例中学习可复用技能的对比框架。每一步中,SSO会在一个未标注批次上运行当前技能,利用部分执行结果生成完整技能探针,并在同一批次上运行这些探针。LLM评判器会比较探针产生的答案、轨迹、产物或终端状态。独立的行为提取器无需查看评判器的决策即可识别行为差异。SSO利用这些决策汇总实例间支持与反对观测行为的证据,随后根据所得证据对行为排序,并从排名最高的行为中生成新的完整技能。仅当新技能在未标注验证集上的表现优于当前技能时,该更新才会被接受。SSO在封闭型和开放型任务上均优于现有无真实标注的提示优化器;在封闭型基准上,它无需任何真实标注反馈,表现接近有时甚至超过最强的基于真实标注的技能优化器。

英文摘要

Agent skills provide frozen large language model (LLM) agents with reusable procedural guidance, and recent work shows that such skills can be optimized with ground-truth (GT) feedback. Many applications, however, lack GT labels, task scores, rewards, or reliable task-specific evaluators. We therefore introduce Self-Supervised Skill Optimization (SSO), a comparative framework that learns a reusable skill from unlabeled task instances alone. At each step, SSO runs the current skill on an unlabeled batch, uses a subset of the resulting executions to generate complete skill probes, and runs the probes on the same batch. An LLM judge compares the resulting answers, trajectories, artifacts, or terminal states. A separate behavior extractor identifies behavioral differences without seeing the judge's decisions. SSO uses these decisions to aggregate evidence for and against the observed behaviors across instances. It then ranks the behaviors by the resulting evidence and renders a new complete skill from the highest-ranked behaviors. The update is accepted only if the new skill outperforms the current one on an unlabeled validation set. SSO outperforms existing GT-free prompt optimizers on both closed-ended and open-ended tasks. On closed-ended benchmarks, it approaches and sometimes exceeds the strongest GT-based skill optimizer without using any GT feedback.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑