arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

State2State:面向大语言模型智能体的环境衍生式中间训练

State2State: Environment-Derived Mid-Training for LLM Agents

Xuanyu Lei, Yiqi Zhu, Chenliang Li, Kaiming Liu, Peng Li, Ming Yan, Jieping Ye, Ya-Qin Zhang, Yang Liu

arXiv 2608.04934首次发表:更新:

发表机构

Institute for AI Industry Research (AIR), Tsinghua University; Institute for AI, Tsinghua University; Institute of Intelligent Computing, Alibaba Group(清华大学人工智能产业研究院; 清华大学人工智能研究院; 阿里巴巴集团智能计算研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出State2State环境衍生式中间训练方法,无需外部任务与监督,可提升LLM智能体性能及学习效率,具备跨环境泛化潜力。

AI 中文摘要

训练大语言模型(LLM)智能体通常依赖于来自专家轨迹的监督微调,或在人工指定任务上借助人工设计的验证器开展的在线强化学习。尽管这两种方法有效,但均受限于外部指定的任务和监督信号,限制了智能体训练的可扩展性与多样性。我们研究一种环境学习范式,其中智能体仅通过与环境交互即可获得交互与操作能力,无需外部指定任务。我们提出State2State,一种环境衍生式中间训练方法,该方法将探索得到的环境状态转化为训练目标,要求智能体到达指定的目标状态。通过从环境探索中衍生任务并基于规则的状态匹配验证成功与否,State2State无需专家监督或手动任务设计即可提供可扩展且可验证的训练目标。在ALFWorld与ScienceWorld上开展的实验表明,State2State作为独立的环境学习阶段,在多数设置下可提升智能体性能;作为下游强化学习的初始化方法,它还能进一步提升最终性能与学习效率,具备跨环境泛化的良好潜力。

英文摘要

Training LLM agents commonly relies on supervised fine-tuning from expert trajectories or online reinforcement learning over human-specified tasks with handcrafted verifiers. Though effective, both remain bottlenecked by externally specified tasks and supervision signals, limiting the scalability and diversity of agent training. We study an environment learning paradigm in which agents acquire interaction and manipulation capabilities solely through environment interaction, without externally specified tasks. We propose State2State, an environment-derived mid-training method that converts explored environment states into training objectives, challenging agents to reach a specified target state. By deriving tasks from environment exploration and verifying success through rule-based state matching, State2State provides scalable and verifiable training objectives without expert supervision or manual task design. Experiments on ALFWorld and ScienceWorld show that State2State improves agent performance as a standalone environment-learning stage in most settings. As initialization for downstream RL, it further improves final performance and learning efficiency, with promising evidence of cross-environment generalization.

CommentsWork in progress

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑