S3Gym:大型语言模型能否将自我测试与自我判断转化为自我提升?
S3Gym: Can LLMs Turn Self-Testing and Self-Judging into Self-Improvement?
浏览论文内容
中文总结 AI 辅助
本研究推出交互式基准S³Gym,评估LLM通过自我测试、自我判断、自我提升能力实现的自我提升,发现其效果依赖任务结构,参数训练等途径各有优劣。
中文摘要 AI 辅助
大型语言模型(LLMs)越来越多地与外部环境交互并积累大量行为经验,但现有的智能体基准大多将其评估为固定策略。因此,目前尚不清楚智能体是否能主动测试自身行为、判断产生的经验,并利用这些经验改进未来的决策。我们推出S³Gym,这是一个用于评估LLM自我提升的交互式基准,涵盖三个耦合能力:自我测试、自我判断和自我提升。S³Gym将宽松探索与严格的保留评估分开,并在七个基于文本的游戏中实现该协议,这些游戏带有可执行的环境验证器。我们评估了三种整合交互经验的途径:直接的历史上下文学习(History ICL)、分数条件总结记忆(score-conditioned Summary Memory)和参数训练(parameter Training)。实验表明,自我提升既非自动也非均匀:上下文级经验可提升部分模型-游戏对的性能,但最有效的途径高度依赖任务结构——当经验可压缩为可复用的战略规则时,总结是有益的;而当成功取决于精确的、依赖状态的信息时,总结的表现往往不如原始历史。参数训练在部分任务上产生显著增益,但也存在改进不稳定及在其他任务上出现严重负迁移的问题。这些发现表明,识别成功的行动是不够的;智能体还必须将反馈转化为可执行且可迁移的策略。S³Gym提供了一个统一框架,用于诊断这一过程并确定阻碍智能体将交互经验转化为可靠自我提升的瓶颈。
英文摘要
Large language models (LLMs) increasingly interact with external environments and accumulate substantial behavioral experience, yet existing agent benchmarks largely evaluate them as fixed policies. It therefore remains unclear whether an agent can actively test its behavior, judge the resulting experience, and use that experience to improve future decisions. We introduce \textbf{S\textsuperscript{3}Gym}, an interactive benchmark for evaluating LLM self-improvement through three coupled capabilities: \textbf{Self-Testing}, \textbf{Self-Judging}, and \textbf{Self-Improvement}. S$^3$Gym separates permissive exploration from strict held-out evaluation and instantiates this protocol in seven text-based games with executable environment verifiers. We evaluate three pathways for incorporating interaction experience: direct History ICL, score-conditioned Summary Memory, and parameter Training. Our experiments reveal that self-improvement is neither automatic nor uniform. Context-level experience improves performance for several model--game pairs, but the most effective pathway depends strongly on the task structure: summaries are beneficial when experience can be compressed into reusable strategic rules, yet often underperform raw history when success depends on precise, state-contingent information. Parameter training produces substantial gains on some tasks, but also exhibits unstable improvement and severe negative transfer on others. These findings show that recognizing successful actions is insufficient; agents must also transform feedback into executable and transferable policies. S$^3$Gym provides a unified framework for diagnosing this process and identifying the bottlenecks that prevent agents from translating interaction experience into reliable self-improvement.
发表机构
- ByteDance Seed(字节跳动Seed)
- M-A-P
机构由 AI 辅助整理,请以论文原文为准。