arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.19072cs.AIcs.CLcs.LG

AI后训练AI缺失了什么?一项实证分析

What is Missing from AI Post-Training AI: An Empirical Analysis

  • Tsinghua University(清华大学)
  • University of Electronic Science and Technology of China(电子科技大学)
  • Renmin University of China(中国人民大学)

机构由 AI 辅助整理,请以论文原文为准。

Joy Jia Yin Lim, Xin Huang, Hao Peng, Yaxi Lu, Xin Cong, Zhong Zhang, Maosong Sun, Yankai Lin

AI总结:

本文通过实证分析发现,LLM智能体后训练时存在策略锁定问题,其缺失的是执行过程中自发重新评估训练策略的机制,而非经验、指导或推理计算。

AI中文摘要:

大型语言模型(LLM)智能体如今可对LLM进行端到端的后训练,它们能编写代码、启动训练、评估检查点并提升下游性能,催生了“AI为AI”的前景。本文指出,这一图景混淆了两种不同能力:执行层面能力,即在选定的训练策略内迭代;策略层面能力,即随着实验证据积累修正高层判断。通过分析大量公开的后训练轨迹语料,我们发现,在不同任务中,智能体的训练策略从一开始就被锁定,剩余全部预算都用于在选定策略内进行局部调整。随后,我们通过逐步升级干预措施,检验三种自然解释——缺失经验、缺失指导、推理不足。大量实验表明:(1)经验驱动的脚手架全面提升了执行性能(在GSM8K上提升12.6个点,在HumanEval上提升40.8个点),但策略保持静态;(2)人类指导能有效重定向初始策略,但训练开始后智能体会退回局部调整循环;(3)额外的推理计算在较简单任务上有收益,但在最难任务上几乎无增益。综上,智能体缺失的既非经验、指导,也非推理计算,而是一种在执行过程中自发重新评估自身策略的机制。

英文摘要:

Large language model (LLM) agents can now post-train an LLM end-to-end, raising the prospect of recursive self-improvement (RSI). Yet this progress is measured by aggregate benchmark scores, which cannot tell whether an agent executes a fixed plan well or strategically revises the plan when it fails. We separate these two capabilities: execution-level capability, iterating within an established training strategy, and strategy-level capability, revising that strategy as experimental evidence accumulates. Analyzing 1,338 post-training trajectories of frontier agents, we find that agents reliably execute post-training but lock into a default strategy, which follows the agent rather than the task, and only 2.1% of transitions between adjacent training runs ever change strategy. We then test whether the agent lacks experience, reasoning, or the decision to switch. (1) Experience improves execution but not the strategy. (2) Additional reasoning compute yields front-loaded gains on easier tasks but refines, rather than revises, the committed strategy. (3) Human review before training changes which strategy the agent locks into, not whether it locks in, whereas a single mid-run instruction outperforms the agent's own continuation by up to 17.44 points under the same budget. In conclusion, what the agent lacks is the decision to reopen a committed strategy and try another one. Realizing RSI therefore calls for interaction protocols and training signals that make strategy revision an explicit, rewarded decision.

↑