管理神经符号强化学习中的动作前提条件:具身智能体的三种放置策略
Managing Action Preconditions in Neuro-Symbolic RL: Three Placement Strategies for Embodied Agents
- Institute of Distributed Intelligent Systems(分布式智能系统研究所)
- University of the Bundeswehr Munich(慕尼黑联邦国防军大学)
- School of Engineering(工程学院)
- The University of Western Australia(西澳大利亚大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究提出将行为知识作为前提贝叶斯网络注入神经符号强化学习,并比较验证器、执行器、学习器三种放置策略,实验表明能显著提升解决方案质量。
AI中文摘要:
人类会将熟悉情境中如何行动的行为知识带入每一项新任务,而不是从头重新学习。强化学习(RL)智能体没有理由不这样做:已知的行为模式无需学习,只需应用。神经符号强化学习通过将符号知识与学习到的策略一起注入,从而将先验知识与强化学习连接起来。知识整合的时机至关重要:选择不当可能产生,例如,幻觉前提条件,这会在智能体于变化环境中行动时表现为安全性和可靠性问题。我们将这种行为知识形式化为一个关于智能体“结构动作”的前提贝叶斯网络(BN)——这些动作的合法性取决于前提条件,例如拾取钥匙、抓取积木、切换门或放下物体。BN限制了这些动作何时可以触发,我们将其注入到强化学习循环中的三个位置:(1)“符号验证器”,仅在推理时咨询,一旦前提条件成立就触发结构动作;(2)“符号执行器”,在训练和推理期间均活跃,在整个学习过程中控制结构动作的使用;(3)“符号学习器”,将知识融入网络,并自行学习结构动作的限制和使用。为了测试这三种变体,我们在两个具有相反机制的基准上进行了实验:一个基于长序列、有序的规划链,另一个基于连续操作。我们在解决方案质量、样本效率和可追溯性方面与强基线进行了比较。收益是显著的。在MiniGrid上,所有三种放置方式都提高了相对于PPO+RND基线的“解决方案质量”,其中符号执行器以98.2%领先,而基线为88.8%。在Fetch上,……
英文摘要:
Humans carry behaviour knowledge of how to act in familiar situations into every new task rather than relearning it from scratch. There is no reason a Reinforcement Learning (RL) agent shouldn't do the same: known behaviour patterns need not be learned, only applied. Neuro-symbolic RL bridges prior knowledge and RL by injecting symbolic knowledge alongside a learned policy. The point at which this knowledge is integrated is critical: a poor choice can produce, for instance, hallucinated preconditions, which surface as safety and reliability problems in agents acting in changing environments. We formalise this behavioural knowledge as a precondition Bayesian network (BN) over the agent's \emph{structural actions} - the actions whose legality depends on preconditions, such as picking up a key, grasping a block, toggling a door, or dropping an object. The BN restricts when these actions may fire, and we inject it into the RL loop at three placements: (1) a \emph{symbolic verifier}, consulted only at inference, that fires a structural action once its preconditions hold; (2) a \emph{symbolic enforcer}, active during both training and inference, that governs structural-action use throughout learning; and (3) a \emph{symbolic learner}, which folds the knowledge into the network and learns the restriction and use of structural actions itself. To test the three variants we run experiments on two benchmarks with opposite regimes: one built on long, ordered planning chains, the other on continuous manipulation. We compare against strong baselines on solution quality, sample efficiency, and traceability. The payoff is substantial. On MiniGrid, all three placements improve the \emph{solution quality} over the PPO+RND baseline, the symbolic enforcer leading at $98.2\%$ against the baseline's $88.8\%$. On Fetch, $\dots$