arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ASPIRE:模型能否从模糊目标中实现自我进化?

Aspire: Can Models Self-Evolve from Vague Goals?

Yuhao Wu, Jingyuan Zhang, Jiajun Shi, Yuxuan Zhang, Xinping Lei, Junting Zhou, Zexuan Wang, Yuchen Wu, Huan Zhou, Duo Wang, Yinzhu Piao, Yongchang Peng, Yunfeng Shi, Jin Chen, Zuo Wang, Jinkai Liu, Jiaheng Liu, Wenxuan Zhang, Shen Yan, Wenhao Huang, Ge Zhang

arXiv 2608.31111首次发表:更新:

发表机构

ByteDance Seed; Singapore University of Technology and Design; M-A-P(字节跳动 Seed; 新加坡科技设计大学; M-A-P)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出ASPIRE基准,探究模型能否从模糊目标实现自我进化,实验发现当前智能体虽能完成基础循环,但权重提升不稳定,最强进化harness仍低于Qwen-Agent基准。

AI 中文摘要

人类的许多重要学习形式都始于模糊目标,比如“成为更优秀的物理学家”或“提升研究能力”。学习者必须解读该目标、识别能力差距、确定学习方式,并判断自身是否真正取得进步。相比之下,现有的大语言模型(LLM)自我进化研究通常以人类指定的任务和评估指标为起点,将自我进化简化为对显式目标的优化,而非决定学习内容与方式。本文提出ASPIRE——一个用于模糊目标驱动自我进化的基准。ASPIRE仅提供自然语言形式的能力目标,下游评估任务则保持隐藏。智能体必须通过选择数据与更新方法、构建训练和验证信号、决定评估时机来将目标具体化。ASPIRE在统一的交互式环境中同时支持模型权重与智能体 harness 的进化,并在专家构建的包含6个目标、共520个条目的隐藏数据集上对生成的系统进行评估。实验结果显示,模糊目标会将搜索努力导向目标解读;当前智能体虽能常规完成训练与 harness 编辑循环,但权重层面的提升仍较为零散且不稳定,最强的进化 harness 仍低于工程化的Qwen-Agent基准。智能体常使用不匹配的数据进行训练,并信任狭隘的自我评估,因此局部提升无法迁移至隐藏评估,持续的搜索与训练还可能抹去先前的改进。

英文摘要

Many important forms of human learning begin with a vague goal, such as "become a better physicist" or "improve at research." Learners must interpret the goal, identify capability gaps, decide how to learn, and determine whether they have actually improved. In contrast, existing work on LLM self-evolution typically begins with tasks and evaluation metrics specified by humans, reducing self-evolution to optimizing an explicit objective rather than deciding what and how to learn. We introduce ASPIRE, a benchmark for vague-goal-driven self-evolution. ASPIRE provides only a natural-language capability goal while downstream evaluation tasks remain hidden. The agent must operationalize the goal by choosing data and update methods, constructing training and validation signals, and deciding when to evaluate. ASPIRE supports both model-weight and agent-harness evolution in a unified interactive environment and evaluates the resulting systems on a hidden, expert-authored set of 520 items spanning six goals. Our experiments show that vague goals redirect search effort toward goal interpretation. Current agents routinely complete training and harness-editing loops, but weight-level gains remain sparse and unstable, and the strongest evolved harness remains below the engineered Qwen-Agent reference. Agents often train on mismatched data and trust narrow self-evaluations, so local gains fail to transfer to hidden evaluation and continued search and training can erase earlier improvements.

Commentshttps://self-developing-agents.github.io/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑