arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从世界反馈中学习:为何模型不确定性在基于模型的强化学习中无法作为风险信号

Learning from World Feedback: Why Model Uncertainty Fails as a Risk Signal in Model-Based RL

Zhaohui Wang

arXiv 2607.16591首次发表:更新:

AI 中文总结

研究探讨 RLxF 中学习信号应源于世界反馈,在安全模型控制中实例化并提炼原则。通过实验表明基于动力学的不确定性惩罚会增加碰撞率,用世界反馈信号可降低碰撞率,提取原则并指出其适用于多种相关方法。

AI 中文摘要

RLxF 计划认为学习信号应来自世界反馈而非内部模型代理。我们在安全的基于模型的控制中实例化这一观点,并提炼出三条具体设计原则。通过实证,在四种跨越 2 倍均方误差范围的世界模型架构中,MPC 规划在统计上等效,基于动力学的不确定性惩罚会使碰撞率从 26%增至 34%。用三种世界反馈信号取代模型内部代理可将碰撞率降至 1 - 14%。机制上,模型不确定性与任务风险的支持空间不同,相关性低。由此提取三条 RLxF 原则,并认为其适用于多种方法。

英文摘要

The RLxF programme argues that learning signals should come from world feedback rather than from internal model proxies. We instantiate this position in safe model-based control and distil it into three concrete design principles. Empirically, across four world-model architectures spanning a 2x MSE range, MPC planning is statistically equivalent (TOST, n=200), and dynamics-based uncertainty penalties increase collision rates from 26% to 34%: the standard MBRL safety proxy is anti-correlated with safety in this regime. Replacing the model-internal proxy with three world-feedback signals (a sensor-derived margin via minimum lidar, a temporal signal via time-to-collision, and an outcome-supervised feedback model g_psi trained on prior collision labels, structurally analogous to outcome-trained reward models in RLHF) reduces collisions to 1-14% without retraining the world model or the planner. The mechanism is structural: model uncertainty has support over state-prediction space, whereas task risk has support over constraint boundaries, with empirical correlation r < 0.15. From this we extract three RLxF principles (ground risk in world outcomes, validate proxies before deployment, and substitute outcome-trained feedback models when direct world signals are unavailable) and argue they apply equally to model-based control and to verifier-based or RLHF approaches in LLM alignment.

CommentsAccepted at the ICML 2026 Workshop on Reinforcement Learning from X Feedback (RLxF). 14 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑