用典范多项式不变量诊断强化学习模拟器与世界模型中的故障
Diagnosing Faults in Reinforcement Learning Simulators and World Models with Canonical Polynomial Invariants
- MIU
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出用典范多项式不变量诊断强化学习模拟器故障,通过筛选与归因程序准确定位违反的物理约束及参数,优于观察空间基线,并应用于多个环境验证其有效性。
AI中文摘要:
大量文献在学习的动力学中构建物理结构,其前提是尊重底层物理的模型预测更好。我们使用从轨迹中恢复并典范化为 $\mathbb{Q}$ 上的约化 Gröbner 基的精确多项式不变量来检验这一前提。在 Acrobot 上,精确性对预测几乎没有益处:一致性正则化器减少了代数残差,但滚动保真度基本不变,而从具有 100% 质量误差的系统中恢复的塑形势能加速学习的效果与正确势能一样有效。精确典范不变量反而被证明对诊断有价值。我们开发了两种程序:筛选(识别被违反的物理约束)和归因(恢复故障不变量并识别负责的物理参数)。为此,我们引入了正规形去flation和商空间恢复。在十五种注入故障中,筛选定位了每一个被破坏的约束且无误报,而观察空间基线无法定位任何故障;归因在所有七种参数故障上恢复了负责的参数。配对差异检验检测到所有故障,表明优势在于定位而非检测。将参考生成器扰动 $10^{-4}$ 可保留 14-15/15 的定位,表明筛选不需要精确性,而理想等式决策能区分仅 $10^{-12}$ 的扰动,表明代数比较需要精确性。应用于十一个 RL 环境的 350 个发布对,诊断未发现模拟器动力学变化的证据,反而揭示了基准实现本身的属性。
英文摘要:
A large literature builds physical structure into learned dynamics on the premise that models respecting the underlying physics predict better. We test that premise using exact polynomial invariants recovered from trajectories and canonicalised as reduced Gröbner bases over $\mathbb{Q}$. On Acrobot, exactness provides little benefit for prediction: a consistency regulariser reduces algebraic residual while leaving rollout fidelity essentially unchanged, and a shaping potential recovered from a system with a 100% mass error accelerates learning as effectively as the correct potential. Exact canonical invariants instead prove valuable for diagnosis. We develop two procedures: screening, which identifies the violated physical constraint, and attribution, which recovers the faulty invariant and identifies the responsible physical parameter. To enable this, we introduce normal-form deflation and quotient-space recovery. Across fifteen injected faults, screening localises every broken constraint with no false alarms, whereas observation-space baselines do not localise any; attribution recovers the responsible parameter on all seven parameter faults. Paired difference tests detect all faults, showing that the advantage is localisation rather than detection. Perturbing reference generators by $10^{-4}$ preserves 14--15/15 localisations, showing that screening does not require exactness, whereas ideal-equality decisions distinguish perturbations of only $10^{-12}$, showing that exactness is required for algebraic comparison. Applied to 350 release pairs across eleven RL environments, the diagnostic finds no evidence of changed simulator dynamics, instead revealing properties of the benchmark implementations themselves.