面向基于神经符号世界模型的零样本任务迁移
Towards Zero-Shot Task Transfer with Neurosymbolic World Models
浏览论文内容
中文总结 AI 辅助
本研究提出一种神经符号世界模型,通过解耦观测重构与奖励预测,实现无需额外环境交互的零样本任务迁移,其泛化能力优于纯神经方法。
中文摘要 AI 辅助
当前最先进的基于模型的强化学习方法会学习神经世界模型,这类模型可通过在潜在空间中规划来改进策略,且无需对底层环境的结构做任何假设。尽管这类模型具有较强的表达能力,但通常依赖于特定任务:它们会学习与训练任务绑定的不可解释的潜在表示,因此难以泛化到新任务。在本研究中,我们提出了一种新颖的世界模型公式,其中奖励预测仅依赖于整个潜在状态中结构化符号组件的子集。将观测重构与奖励预测解耦,使我们能够学习可实现零样本适应的世界模型,即无需进一步的环境交互,就能适应在相同符号状态空间上定义的新奖励函数。我们讨论了学习这些神经符号世界模型的主要优势和挑战,并展示了我们的方法相较于纯神经方法具有更强的泛化特性。
英文摘要
State-of-the-art model-based reinforcement learning methods learn neural world models that allow policy improvement by planning in a latent space, without assumptions on the structure of the underlying environment. While expressive, these models are generally task-dependent: they learn uninterpretable latent representations that are tied to the training task and thus hard to generalize to new tasks. In this work, we present a novel world model formulation where the reward prediction only depends on a subset of structured, symbolic components of the whole latent state. Decoupling observation reconstruction and reward prediction allows us to learn world models that can adapt zero-shot, i.e. without further environment interactions, to new reward functions defined over the same symbolic state space. We discuss the main advantages and challenges of learning these neurosymbolic world models and demonstrate the strong generalisation properties of our approach over purely neural methods.