发表机构
University of Science and Technology of China; Shanghai AI Laboratory; MIT(中国科学技术大学; 上海人工智能实验室; 麻省理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出Belayer系统,针对LLM智能体强化学习训练的生成引擎和环境故障设计容错机制,可降低无故障训练开销,大幅缩短工作节点与环境故障的恢复时间。
AI 中文摘要
大型语言模型(LLM)智能体越来越多地在长时序、沙箱环境中通过强化学习进行训练。与传统强化学习不同,智能体强化学习将GPU密集型的生成引擎与有状态的环境容器相结合,这些容器的操作可能产生可见的副作用,例如文件编辑、命令执行和依赖项安装。单个轨迹可跨越多轮生成与环境交互,因此组件故障可能会丢弃已完成的工作,或使模型暴露于与其上下文不一致的环境状态中。然而,现有系统缺乏针对这种分布式执行模型的高效且正确的恢复机制。本文提出了Belayer,一种面向LLM智能体强化学习训练的高效容错系统。Belayer可处理生成引擎和环境执行中的故障,同时以低无故障开销为目标。针对限定范围的工作节点本地生成故障,Belayer为每个预初始化的影子工作节点配备了选择性GPU状态复用协议,该协议在所有者和GPU健康检查后保留独立拥有的权重和原始KV Arena分配,重新初始化工作节点本地状态,并从记录的token前缀重建请求特定的KV内容。针对环境故障,Belayer引入了完整检查点和完整恢复机制,以联合捕获和恢复容器文件系统与运行时状态,并协调恢复后的环境与LLM上下文,以保持前缀一致性。自适应策略会在预测间隔足够长时,适时将全状态检查点与自然的LLM推理间隙重叠。实验结果显示,无故障训练期间的测量开销较低,工作节点恢复时间相比完整引擎冷启动最多可缩短42倍,环境故障恢复速度则提升1.5至3.5倍。
英文摘要
Large language model (LLM) agents are increasingly trained with reinforcement learning in long-horizon, sandboxed environments. Unlike conventional RL, agentic RL couples GPU-intensive rollout engines with stateful environment containers whose actions may produce visible side effects, such as file edits, command execution, and dependency installation. A single trajectory can span many rounds of gen- eration and environment interaction, so a component failure can discard completed work or expose the model to an environment state that is inconsistent with its context. However, existing systems lack efficient and correct recovery mechanisms for this distributed execution model. This paper presents Belayer, an efficient fault-tolerant system for LLM agentic RL training. Belayer handles failures in both rollout engines and environment execution while targeting low failure-free overhead. For scoped worker-local rollout failures, Belayer equips each pre-initialized shadow worker with a selective GPU-state reuse protocol that retains independently owned weights and raw KV-arena allocations after owner and GPU health checks, reinitializes worker-local state, and rebuilds request-specific KV contents from logged token prefixes. For environment failures, Belayer introduces full checkpoint and full restore to jointly capture and restore container file-system and runtime state, and coordinates the recovered environment with the LLM context to preserve prefix consistency. An adaptive policy opportunistically overlaps full-state checkpointing with natural LLM inference bubbles when the predicted interval is long enough. Empirical results show low measured overhead during failure-free training, a worker-recovery-time reduction of up to 42 times faster compared with a full engine cold start, and 1.5 to 3.5 times faster recovery from environment failures.