发表机构
University of Michigan(密歇根大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对机器人观察中物理变量交互作用问题,提出PRISM策略表示,通过因式分解多项式模块展现高阶交互特征,在强化和模仿学习中应用,提升了人形机器人运动等性能,还能产生无传感器柔顺行为,表明多项式表示应成运动控制标准选择。
AI 中文摘要
机器人策略通常是将观察映射到动作的多层感知器(MLP)。然而,机器人观察是物理变量,许多与动作相关的线索并非来自单个变量,而是来自它们的交互作用;功率、惯性效应、接触、滑动和柔顺性取决于可观察信号之间的乘积。我们引入了PRISM,一种能使可观察物理变量之间的多项式交互作用变得明确、可学习且紧凑的策略表示。PRISM使用因式分解多项式模块来高效地展现高阶交互特征。在强化学习中,它保留标准MLP主干,并在其后应用逐元素多项式函数且逐步激活。在模仿学习中,它用端到端训练的多项式层取代扩散策略中的线性本体感受调节。在人形机器人运动和丰富接触的操作中,PRISM比标准MLP策略和具有匹配容量的更大MLP性能更好,表明交互结构不能仅由容量替代。它还能在无力、扳手、触觉输入、接触标签或导纳控制的情况下产生无传感器的柔顺行为。这些结果表明多项式表示应成为具身运动控制的标准架构选择。
英文摘要
Robot policies are typically MLPs mapping observations to actions. Yet robot observations are physical variables, and many action-relevant cues arise not from individual variables but from their interactions; power, inertial effects, contact, slip, and compliance depend on products among observable signals. We introduce PRISM, a policy representation that makes polynomial interactions among observable physical variables explicit, learnable, and compact. Rather than listing all polynomial terms, PRISM uses a factorized polynomial module to expose higher-order interaction features efficiently. In reinforcement learning, it keeps the standard MLP backbone but applies a gradually activated element-wise polynomial function after it. In imitation learning, it replaces linear proprioceptive conditioning in Diffusion Policy with a polynomial layer trained end-to-end. Across humanoid locomotion and contact-rich manipulation, PRISM improves performance over standard MLP policies and larger MLPs with matched capacity, showing that interaction structure cannot be replaced by capacity alone. It also yields sensorless compliant behavior without force, wrench, tactile input, contact labels, or admittance control. These results suggest that polynomial representations should become a standard architectural choice for embodied motor control. The project page is available at https://lsh3163.github.io/prism/