arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

为强化学习指定奖励函数而无需环境采样

Specifying Reward Functions for RL Without Environment Sampling

Stephane Hatgis-Kessell, W. Bradley Knox, Emma Brunskill

arXiv 2609.15544首次发表:更新:

发表机构

Stanford University; The University of Texas at Austin(斯坦福大学; 德克萨斯大学奥斯汀分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出EARS方法,利用LLM构建特征并采样想象轨迹,从偏好中学习奖励函数,无需环境交互,在三个长时域任务中优于直接提示LLM的基线。

AI 中文摘要

让人类利益相关者能够指定导致其期望结果的奖励函数,是部署强化学习智能体的一个关键挑战。基于偏好的方法,如在线RLHF,可以减轻手动奖励设计的负担,但它们需要反复训练策略、从真实世界采样轨迹并征求反馈,这使得它们在环境交互计算成本高昂或不安全的环境中不切实际。我们引入了无经验自主奖励指定(EARS),一种无需环境交互即可从偏好中学习奖励函数的方法。我们的方法使用一个结构化的LLM中介过程,从任务描述和环境观测空间构建一小部分富有表现力的奖励特征,然后在该特征空间中策略性地采样想象轨迹,并从对想象轨迹对的偏好中学习特征权重。我们在三个长时域领域进行评估:流行病封锁法规设计、糖尿病患者的胰岛素给药以及高速公路上的自动驾驶车辆控制。我们将EARS与同样能够在无环境交互情况下指定奖励的基线方法进行比较——即直接提示LLM生成奖励函数的方法。当从真实偏好标签或由LLM标记的偏好中学习时,EARS设计的奖励函数比这些基线更符合产生偏好或LLM上下文的真实奖励函数。这些结果表明,基于偏好的奖励指定在没有环境采样的情况下仍然有效,从而在收集真实轨迹成本高昂或不可行的环境中实现实用的奖励设计。

英文摘要

Enabling human stakeholders to specify reward functions that lead to their desired outcomes is a key challenge in deploying reinforcement learning agents. Preference-based methods such as online RLHF can reduce the burden of manual reward design, but they require repeatedly training policies, sampling trajectories from the real world, and eliciting feedback, making them impractical in settings where environment interaction is computationally expensive or unsafe. We introduce Experience-Free Autonomous Reward Specification (EARS), a method for learning reward functions from preferences without environment interaction. Our approach uses a structured LLM-mediated process to construct a small set of expressive reward features from a task description and the environment observation space, then strategically samples imagined trajectories in this feature space and learns feature weights from preferences over the imagined trajectory pairs. We evaluate on three long-horizon domains: pandemic lockdown regulation design, insulin administration for diabetes patients, and autonomous vehicle control on a highway. We compare EARS to baselines that also enable reward specification without environment interaction--namely, methods that directly prompt an LLM to generate a reward function. When learning from either ground-truth preference labels or preferences labeled by a LLM, EARS designs reward functions that are more aligned with the ground truth reward function that produced the preferences or LLM context than these baselines. These results suggest that preference-based reward specification remains effective without environment sampling, enabling practical reward design in settings where collecting real trajectories is costly or infeasible.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑