arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

学习终端智能体的可泛化行为

Learning Generalizable Behaviors for Terminal Agents

Yihang Yao, Bo Pang, Xuan Phi Nguyen, Ding Zhao, Shafiq Joty, Semih Yavuz

arXiv 2608.22631首次发表:更新:

发表机构

Salesforce AI Research; Carnegie Mellon University(Salesforce人工智能研究院; 卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对终端智能体泛化问题,提出智能体组合泛化假设,开发训练方案River,其用不足30%的TMax环境使多规模模型在两个基准上RL增益平均提升超100%,且性能优于开源8B模型、可跨多维度泛化。

AI 中文摘要

终端智能体是大型语言模型(LLMs)的一个极具吸引力的应用,具备深度融入用户日常工作流程的潜力。强化学习(RL)是提升其能力的关键技术,这使得可扩展训练环境成为核心挑战。由于公开的真实用户交互数据稀缺,合成环境提供了一种实用替代方案,但通常存在领域差距和保真度有限的问题,导致泛化能力较差。现有工作主要在扩大合成环境的数量和多样性,而奖励信号质量和决定泛化的机制仍未得到充分探索。我们研究了RL如何改进终端智能体,并提出了智能体组合泛化假设:RL并非从头教授新的特定领域技能,而是主要塑造高级决策行为,这些行为会组合和路由预训练及监督微调(SFT)期间获得的低级技能。这一解释与我们的实证结果一致,且表明决定哪些行为会被强化的验证器质量,比单纯增加环境数量或多样性更为重要。受这一见解启发,我们提出了River,这是一种简单的训练方案,通过过滤低质量环境并使用过程级行为正则化增强结果奖励来提升奖励质量。使用该方案,我们经RL训练的智能体在四个终端智能体基准测试中,在评估的开源RL训练8B模型中取得了最佳性能。River还可跨模型家族、规模、智能体 harness 和 RL 目标进行泛化。使用少于30%的TMax训练环境,River使2B至27B模型在Terminal-Bench-Lite和Terminal-Bench-v2.1上的RL增益分别平均提升了106%和30%。

英文摘要

Terminal agents are a compelling application of large language models (LLMs), with the potential to integrate deeply into users' daily workflows. Reinforcement learning (RL) is a key technique for improving their capabilities, making scalable training environments a central challenge. Since public real-user interaction data are scarce, synthetic environments provide a practical alternative, but often suffer from domain gaps and limited fidelity, leading to poor generalization. Existing work mainly scales the quantity and diversity of synthetic environments, while reward-signal quality and the mechanisms governing generalization remain under-explored. We study how RL improves terminal agents and propose the Agentic Compositional Generalization hypothesis: rather than teaching new domain-specific skills from scratch, RL primarily shapes high-level decision-making behaviors that compose and route low-level skills acquired during pre-training and supervised fine-tuning (SFT). This account is consistent with our empirical results and suggests that verifier quality, which determines which behaviors are reinforced, is more important than simply increasing environment quantity or diversity. Motivated by this insight, we propose River, a simple training recipe that improves reward quality by filtering low-quality environments and augmenting outcome rewards with process-level behavior regularization. Using this recipe, our RL-trained agent achieves the best performance among evaluated open-source RL-trained 8B models across four terminal-agent benchmarks. River also generalizes across model families, scales, agent harnesses, and RL objectives. Using fewer than 30% of the TMax training environments, River improves RL gains by 106% and 30% on average for models ranging from 2B to 27B on Terminal-Bench-Lite and Terminal-Bench-v2.1, respectively.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑