arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过动力学嵌入从无动作时间序列学习可迁移策略

Learning Transferable Policies from Action-free Time Series Through Dynamical Embeddings

Niklas Emonds, Georgia Koppe

arXiv 2610.03065首次发表:更新:

发表机构

Hector Institute for AI in Psychiatry (HITKIP); Central Institute of Mental Health (CIMH); Heidelberg University; Interdisciplinary Center for Scientific Computing (IWR); Hertie Institute for AI in Brain Health; University of Tübingen(赫克托人工智能精神病学研究所; 中央精神健康研究所; 海德堡大学; 跨学科科学计算中心; 赫蒂人工智能脑健康研究所; 蒂宾根大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出层次化模型强化学习框架,利用低维动力学嵌入共享结构,从无动作记录学习可迁移控制策略,在Lorenz-63和双摆系统上验证了提升的迁移性能与机制分析能力。

AI 中文摘要

从无动作记录中学习控制具有挑战性,因为干预效应未被观测到,且策略可能利用重建动力学中的误差。我们提出了一种基于模型的层次化强化学习框架,利用相关系统间的共享结构,从无动作记录中学习系统特定的控制策略。层次化动力学系统重建模型通过低维嵌入捕获共享动力学和个体差异。这些嵌入随后被复用来参数化共享的策略网络和价值网络,将重建动力学的差异与控制差异联系起来。策略完全在显式干预模型(含加性潜在扰动)下的仿真中训练。分段线性循环神经网络支持对受控动力学进行机制分析,而基于解码器的约束使干预的即时效应在观测空间中可解释,并允许对一种模态进行干预,同时保护另一种模态免受直接操纵。在Lorenz-63和双摆系统上,层次化策略相比独立训练的策略提升了迁移效果。在Lorenz-63上,它们还比使用相同重建模型的重复规划获得更高的平均奖励,与使用受控交互训练的方法表现相当,并且仅经过嵌入推理即可泛化到策略训练中未出现的系统。对神经-行为记录的应用展示了在约束神经扰动下对预测运动的抑制。这些发现共同表明,共享的动力学表示支持从无动作记录中进行可迁移控制和机制性假设生成。

英文摘要

Learning control from action-free recordings is challenging because intervention effects are unobserved and policies may exploit errors in reconstructed dynamics. We present a hierarchical model-based reinforcement learning framework that uses shared structure across related systems to learn system-specific control policies from action-free recordings. A hierarchical dynamical system reconstruction model captures shared dynamics and individual variation through low-dimensional embeddings. These embeddings are then reused to parameterize shared policy and value networks, linking differences in reconstructed dynamics to differences in control. Policies are trained entirely via simulation under an explicit intervention model with additive latent perturbations. Piecewise-linear recurrent neural networks enable mechanistic analyses of the controlled dynamics, while decoder-based constraints make the immediate effects of interventions interpretable in observation space and permit interventions on one modality while protecting another from direct manipulation. On Lorenz-63 and double-pendulum systems, hierarchical policies improve transfer over independently trained policies. On Lorenz-63, they also achieve a higher mean reward than repeated planning with the same reconstructed models, perform comparably to methods trained with controlled interactions, and generalize to systems absent from policy training after embedding inference alone. Applications to neural-behavioral recordings demonstrate suppression of predicted movement under constrained neural perturbations. Together, these findings show how shared dynamical representations support transferable control and mechanistic hypothesis generation from action-free recordings.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑