arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.05473cs.AI

面向语言模型智能体的具有稳定时间抽象的分层强化学习

Hierarchical Reinforcement Learning with Stable Temporal Abstraction for Language Model Agents

Shayan Mohajer Hamidi, Yize Cheng, Yuanda Xu, Zhengze Zhou, Alborz Geramifard

首次发表
浏览论文内容

中文总结 AI 辅助

提出STAC方法,通过约束边界策略优化解决分层强化学习智能体的时间抽象不稳定性,在ALFWorld和WebShop上显著提升成功率。

中文摘要 AI 辅助

分层强化学习通过围绕持久子目标组织原始动作,并在多个时间尺度上分配信用,从而改善长时程控制。最近的分层语言智能体通过明确分离子目标规划与动作执行,将这些优势带入交互式任务。然而,我们观察到,显式分层本身并不能决定所产生的时间抽象的稳定性:学习到的边界策略可能几乎每一轮都替换子目标,使其实际上成为瞬时的,或者在子目标已不再适用后仍保留它。我们称这种现象为时间抽象不稳定性。我们提出通过约束优化的稳定时间抽象(STAC),一种约束边界策略优化方法,将过早重新规划和过时保留表示为约束成本。STAC仅将所得的拉格朗日成本应用于采样的边界决策,而保持底层算法的奖励、评论家目标、子目标优势和原始动作优势不变。在两种主干和两个基准上,STAC在ALFWorld和WebShop上使用Qwen3-0.6B时,相对于强分层基线分别提高了8.1和7.9个百分点的成功率,使用Llama-3.2-1B-Instruct时分别提高了23.5和15.8个百分点。

英文摘要

Hierarchical reinforcement learning improves long-horizon control by organizing primitive actions around persistent subgoals and assigning credit at multiple temporal scales. Recent hierarchical language agents bring these benefits to interactive tasks by explicitly separating subgoal planning from action execution. We observe, however, that an explicit hierarchy does not by itself determine how stable the resulting temporal abstraction is: the learned boundary policy may replace the subgoal almost every turn, making it effectively transient, or retain a subgoal after it has stopped being appropriate. We call this temporal abstraction instability. We propose Stable Temporal Abstraction via Constrained Optimization (STAC), a constrained boundary-policy optimization method that represents premature replanning and stale persistence as constraint costs. STAC applies the resulting Lagrangian costs only to the sampled boundary decision, leaving the underlying algorithm's rewards, critic targets, subgoal advantages, and primitive-action advantages unchanged. Across two backbones and two benchmarks, STAC improves success over a strong hierarchical baseline by $8.1$ and $7.9$ points on ALFWorld and WebShop with Qwen3-0.6B, and by $23.5$ and $15.8$ points with Llama-3.2-1B-Instruct.

发表机构

  • LinkedIn Corporation(领英公司)
  • University of Maryland(马里兰大学)

机构由 AI 辅助整理,请以论文原文为准。

↑