Agent-G$^2$:面向智能体强化学习的高斯引导框架
Agent-G$^2$: Gaussian Guidance for Agentic Reinforcement Learning
浏览论文内容
中文总结 AI 辅助
Agent-G$^2$ 是面向智能体强化学习的高斯引导框架,通过在线估计高斯分布参数确定引导深度,在 ALFWorld 等任务上以更低成本实现优于基线的性能。
中文摘要 AI 辅助
基于提示的强化学习通过在每次策略 rollout 前保留专家轨迹的前缀,让策略从更接近成功的状态开始探索,从而解决长程智能体任务中的奖励稀疏问题。其有效性取决于引导深度,即保留轨迹的长度。现有方法将该深度视为确定性标量:调度方法在所有样本中共享一个值,忽略任务间异质性;逐样本探测方法单独估计深度,但需额外的 rollout 成本。我们发现,有用的引导深度分布在一个深度区间内,其信息性在区间中心附近近似呈高斯分布,而非集中于单一最优值。我们提出 Agent-G$^2$,一种高斯引导框架,它从高斯分布中抽取每个任务的深度,其中心和扩散度可从已收集的用于策略优化的 rollout 中在线估计,无需探测 rollout 或学习深度预测器。中心结合了全局基线与按簇划分的难度,扩散度跟踪簇内方差。我们在 ALFWorld 和 WebShop 上使用 Qwen2.5-1.5B / 7B-Instruct 评估 Agent-G$^2$。Agent-G$^2$ 在 ALFWorld 上比最强的基于提示、无提示和 Aux-RL 基线分别高出 2.3、3.9 和 7.4 个点,且成本仅为逐样本探测的三分之一以下。
英文摘要
Hint-based reinforcement learning addresses reward sparsity in long-horizon agentic tasks by retaining a prefix of an expert trajectory before each rollout, letting the policy explore from a state closer to success. Its effectiveness hinges on the guidance depth: how much of the trajectory to keep. Existing methods treat this depth as a deterministic scalar. Scheduled approaches share one value across samples and ignore per-task heterogeneity; per-sample probing estimates it separately at the cost of extra rollouts. We find that useful guidance occupies a band of depths whose informativeness profile is approximately Gaussian around the band center, rather than concentrating at a single optimal point. We propose Agent-G$^2$, a Gaussian guidance framework that draws the depth per task from a Gaussian whose center and spread are estimated online from rollouts already collected for policy optimization, requiring no probe rollouts or learned depth predictor. The center combines a global baseline with per-cluster difficulty, and the spread tracks within-cluster variance. We evaluate Agent-G$^2$ on ALFWorld and WebShop on Qwen2.5-1.5B / 7B-Instruct. Agent-G$^2$ outperforms the strongest hint-based, hint-free, and Aux-RL baselines on ALFWorld by 2.3 / 3.9 / 7.4 points at under one-third the rollout cost of per-sample probing.
发表机构
- Baidu Inc.(百度公司)
- Shandong University(山东大学)
- Zhejiang University(浙江大学)
机构由 AI 辅助整理,请以论文原文为准。