不要遮蔽环境:观察监督如何改变智能体在强化学习下的探索
Don't Mask the Environment: Observation Supervision Changes How Agents Explore Under RL
浏览论文内容
中文总结 AI 辅助
针对标准SFT仅监督动作而忽略观察的问题,提出ActObs联合监督观察令牌,在不增加开销下提升GRPO后的探索与pass@k,并保留策略初始化。
中文摘要 AI 辅助
智能体轨迹记录了智能体的行为及其后续结果。然而,标准的监督微调(SFT)仅对智能体生成的动作令牌施加损失,将环境观察作为上下文而非预测目标。我们质疑这一惯例是否为后续的强化学习提供了最佳初始化。我们引入了ActObs,该方法同时监督每条轨迹中已存在的观察令牌。尽管部署的智能体从不生成观察,学习预测它们能鼓励策略在不增加数据、参数、序列令牌或前向传播的情况下对动作后果进行建模。在SFT之后,两种方法表现相似,但在GRPO之后出现分歧。在Qwen3-4B上,基于ActObs的GRPO在Terminal-Bench 2.0的每个评估采样预算下,其pass@k均高于仅动作对应方法。在Qwen3-8B上,它牺牲了一些pass@1的可靠性,以换取更高的pass@k(在pass@16时+3.4个百分点),并解决了更多不同的任务。该优势扩展到aider-polyglot上的跨领域代码编辑(在4B时pass@1+4.2个百分点),这些任务在SFT和RL期间均未见过。ActObs在RL期间保留了更多熵,同时需要更少的策略移动,使最终策略更接近其SFT初始化。我们的分析将这种差异追溯到SFT:动作和观察梯度迅速变得正交,而仅动作训练留下较大的残余观察梯度,并使环境预测性能低于基础模型。联合监督防止了这种单侧特化,保留了后果预测能力,并为下游探索做好了策略准备。
英文摘要
Agent trajectories record what an agent does and what happens next. Yet standard supervised fine-tuning (SFT) applies loss only to agent-authored action tokens, using environment observations as context but not as prediction targets. We ask whether this convention provides the best initialization for subsequent reinforcement learning. We introduce ActObs, which also supervises the observation tokens already present in each trajectory. Although deployed agents never generate observations, learning to predict them encourages the policy to model action consequences without adding data, parameters, sequence tokens, or forward passes. The methods perform similarly after SFT but diverge after GRPO. On Qwen3-4B, GRPO from ActObs achieves higher pass@k at every evaluated sampling budget than its action-only counterpart on Terminal-Bench 2.0. On Qwen3-8B, it trades some pass@1 reliability for higher pass@k (+3.4 pp at pass@16) and solves more distinct tasks. The advantage extends to cross-domain code editing on aider-polyglot (+4.2 pp at pass@1 at 4B), whose tasks are unseen during SFT and RL. ActObs retains more entropy during RL while requiring less policy movement, leaving the final policy closer to its SFT initialization. Our analysis traces this difference to SFT: action and observation gradients rapidly become orthogonal, while action-only training leaves a large residual observation gradient and degrades environment prediction below the base model. Joint supervision prevents this one-sided specialization, preserving consequence prediction and preparing the policy for downstream exploration.
发表机构
- University of Maryland(马里兰大学)
- AWS AI Labs(亚马逊云科技人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。