向量符号策略梯度
Vector Symbolic Policy Gradient
浏览论文内容
中文总结 AI 辅助
该研究提出VSPG算法,以单位范数超向量表示离散动作,通过优势加权超向量捆绑更新动作,可实现样本高效学习且推理内存不增加,还证明其动作选择在随机比特翻转下稳定,连接了三类方法并提供鲁棒性保证。
中文摘要 AI 辅助
我们用向量符号策略梯度(Vector-Symbolic Policy Gradient, VSPG)回答该问题,这是一种离散动作的Actor,通过单位范数超向量表示每个动作,并根据其与编码状态的相似度对动作打分。在标准softmax策略梯度替代函数下,我们证明其更新过程恰好是优势加权超向量捆绑后归一化,因此支持标准优势估计器。我们进一步表明,每个训练后的动作超向量是固定大小的压缩核记忆,存储着访问状态上的优势加权核展开,并根据编码器诱导的相似度传递证据,这提供了一种具体机制,可在不增加推理时内存的情况下支持样本高效学习。最后,对于双极动作记忆,我们证明在随机比特翻转下贪心动作选择是稳定的,失败概率随超向量维度呈指数衰减。因此,VSPG连接了VSA动作记忆、对数线性策略梯度和核策略搜索,同时提供了定量鲁棒性保证。
英文摘要
We answer this question with Vector-Symbolic Policy Gradient (VSPG), a discrete-action actor that represents each action by a unit-norm hypervector and scores it by similarity to the encoded state. Under the standard softmax policy-gradient surrogate, we prove that its update is exactly advantage-weighted hypervector bundling followed by normalization, and therefore supports standard advantage estimators. We further show that each trained action hypervector is a fixed-size compressed kernel memory, storing an advantage-weighted kernel expansion over visited states and transferring evidence according to the encoder-induced similarity. This provides a concrete mechanism that can support sample-efficient learning without increasing inference-time memory. Finally, for bipolar action memories, we prove that greedy action selection is stable under random bit flips, with failure probability decaying exponentially in the hypervector dimension. VSPG thus connects VSA action memories, log-linear policy gradients, and kernel policy search while providing a quantitative robustness guarantee.
发表机构
- University of California, Irvine(加州大学欧文分校)
- Intel Corporation(英特尔公司)
- Johns Hopkins University(约翰斯·霍普金斯大学)
- Northeastern University(东北大学)
机构由 AI 辅助整理,请以论文原文为准。