发表机构
Central South University(中南大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对模拟器到目标偏移下的模型选择问题,提出MC-WM方法,通过数据分区和校准风险路由选择模型,并在MuJoCo偏移上验证了其有效性。
AI 中文摘要
基于模型的强化学习(MBRL)可以利用模拟经验,但模拟器到目标的偏移会产生模型选择问题:在目标数据有限的情况下,直接校正模拟器和直接拟合目标都可能失败。我们引入了模型校正世界模型(MC-WM),该方法将初始目标数据划分为互斥的拟合、选择和校准分区,并部署具有较低标准化校准风险的模型族。学习到的置信度信号和确定性有效性谓词对一步想象的策略更新进行加权,而无需重写物理奖励。我们在三个受控的多关节动力学与接触(MuJoCo)偏移上评估了540个独特的报告运行单元;一个精确路由单元在部署前工件门控后被重复,共完成541次执行。
英文摘要
Model-based reinforcement learning (MBRL) can exploit simulated experience, but a simulator-to-target shift creates a model-selection problem: correcting the simulator and fitting the target directly can each fail under limited target data. We introduce the Model-Corrected World Model (MC-WM), which separates initial target data into disjoint fit, selection, and calibration partitions and deploys the family with lower standardized calibration risk. A learned confidence signal and deterministic validity predicates weight one-step imagined policy updates without rewriting physical rewards. We evaluate 540 unique reported run cells across three controlled Multi-Joint dynamics with Contact (MuJoCo) shifts; one exact-routing cell was repeated after a pre-deployment artifact gate, giving 541 completed executions.