发表机构
Carnegie Mellon University(卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对部分可观测机器人操作中状态混叠导致模仿学习性能差的问题,提出用贝叶斯信念结构化表示隐藏状态来条件化扩散策略,在两个领域显著优于基线并接近特权性能。
AI 中文摘要
机器人操作任务通常涉及无法直接观察、必须通过顺序物理交互来推断的隐藏状态信息。在这种部分可观测的环境中,直接将模仿学习策略基于最近的原始观测历史进行条件化会导致性能不佳。这是由于状态混叠现象,即相同的观测可能源自不同的隐藏状态,而策略对于相同输入接收到相互冲突的动作标签。为了实现具有历史感知的去混叠能力,我们提出使用贝叶斯信念将扩散策略条件化于隐藏状态的结构化表示上,该表示同时暴露当前最可能的状态估计和剩余不确定性。这种表示用结构化、紧凑的输入替代原始历史,使策略能够基于信念不确定性在探索性和利用性行为之间隐式调节,而无需显式的模式切换或奖励塑形。我们在两个具有定性不同信念表示的领域进行评估:通过触觉传感进行玉米秆夹持器对齐的连续信念,以及用于门闩开启的离散分类分布。在这两个领域中,信念条件化策略显著优于仅观测的基线,并接近特权地面真值性能,消融实验表明策略在推理时根据信念不确定性调整其探索行为。
英文摘要
Robot manipulation tasks often involve hidden state information that cannot be directly observed and must be inferred through sequential physical interactions. In such partially observable settings, conditioning an imitation learning policy directly on the recent raw observation history leads to poor performance. This is due to state aliasing, wherein identical observations may arise from different hidden states, and the policy receives conflicting action labels for the same input. To enable history-aware disambiguation capability, we propose conditioning a diffusion policy on a structured representation of the hidden states using Bayesian belief that exposes both the current most likely state estimate and the remaining uncertainty. This representation replaces raw history with a structured, compact input, enabling the policy to implicitly modulate between exploratory and exploitative behaviors based on belief uncertainty, without explicit mode switching or reward shaping. We evaluate across two domains with qualitatively different belief representations: a continuous belief for cornstalk gripper alignment via tactile sensing, and a discrete categorical distribution for latched door opening. In both domains, the belief-conditioned policy substantially outperforms the observation-only baseline and approaches privileged ground-truth performance, with ablations illustrating that the policy adapts its exploration behavior depending on the belief uncertainty at inference time.