arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

原位随机伴随训练的物理策略梯度定理

Physical policy gradient theorem for in situ stochastic-adjoint training

William Tuxbury, Zin Lin

arXiv 2609.05808首次发表:更新:

发表机构

Virginia Tech(弗吉尼亚理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出物理策略梯度定理的随机伴随梯度估计器,通过非简并扩散解除互易性限制,并仅用测量轨迹训练非线性谐振子网络对抗时间调制。

AI 中文摘要

原位伴随训练直接从测量中提取参数梯度,但迄今仅限于互易或受限系统。在此,我们引入策略梯度定理的物理对应物:一种随机伴随梯度估计器,通过以非简并扩散换取互易性来解除这些限制。作为验证,我们训练了一个非线性谐振子网络,其自身动力学提供策略,仅凭测量的随机轨迹梯度对抗时间调制,无需有限差分或单独的伴随实验。

英文摘要

In situ adjoint training extracts parameter gradients directly from measurement, but has so far been limited to reciprocal or restricted systems. Here, we introduce the physical counterpart of the policy gradient theorem: a stochastic-adjoint gradient estimator that lifts these constraints by trading reciprocity for nondegenerate diffusion. As validation, we train a nonlinear resonator network, whose own dynamics supply the policy, against antagonistic temporal modulations with gradients from measured stochastic trajectories alone, without finite differences or a separate adjoint experiment.

Comments9 pages, 2 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑