arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.14180cs.LGcs.AI

RENEW:迈向从偏好中学习世界模型并修复模型利用问题

RENEW: Towards Learning World Models and Repairing Model Exploitation from Preferences

Logan Mondal Bhamidipaty, Mykel Kochenderfer, Subramanian Ramamoorthy

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对世界模型在数据覆盖薄弱时易受模型利用问题,提出从人类反馈中学习动力学(DLHF),但朴素DLHF样本效率低,于是引入RENEW利用认知不确定性聚焦微调,经实验评估,RENEW提高样本效率等,为解决离线模型强化学习利用问题提供新途径。

中文摘要 AI 辅助

世界模型在离线强化学习中广泛应用以提高样本效率并生成固定数据集之外的经验。但它们在数据覆盖薄弱时易受模型利用问题影响。先前工作要么通过收集更多专家示范(通常昂贵、不安全或不可用),要么通过避免不确定区域的保守算法(限制泛化)来解决。我们提出直接利用人类对想象中的展开的偏好来修复利用问题,将其形式化为从人类反馈中学习动力学(DLHF),这是一种基于学习到的动力学模型下轨迹对数似然的Bradley - Terry偏好损失。然而,朴素的DLHF样本效率低,所以我们引入RENEW,它利用认知不确定性聚焦于模型最易被利用之处的微调。我们在多个环境中评估发现,朴素的DLHF需要大量偏好预算,而RENEW通过提高样本效率、限制灾难性遗忘和减少预训练世界模型中的利用问题,使该框架变得实用。总之,我们的结果初步证明偏好可直接监督世界模型动力学,为解决基于离线模型的强化学习中的利用问题提供了新方法。

英文摘要

World models are widely used in offline reinforcement learning (RL) to improve sample efficiency and generate experience beyond a fixed dataset. However, they are vulnerable to model exploitation where data coverage is thin. Prior work addresses this either by collecting more expert demonstrations, which is often expensive, unsafe, or unavailable, or by conservative algorithms that avoid uncertain regions, which limits generalization. We propose instead to repair exploitation directly using human preferences over imagined rollouts, leveraging the strong intuitive physics that allows humans to easily spot egregious dynamics hallucinations. We formalize this as Dynamics Learning from Human Feedback (DLHF), a Bradley-Terry preference loss over trajectory log-likelihoods under a learned dynamics model. Unfortunately, naive DLHF is sample inefficient, so we introduce RENEW, which uses epistemic uncertainty to focus finetuning where the model is most exploitable. We evaluate on several Jumanji and classic control environments and find that while naive DLHF requires an outsize preference budget, RENEW makes the framework practical by improving sample efficiency, limiting catastrophic forgetting, and reducing exploitation in pretrained world models. Taken together, our results provide initial evidence that preferences can supervise world model dynamics directly, offering a new approach to addressing exploitation in offline model-based RL.

↑