arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.30500cs.LGcs.AI

PolicyAttention:软最大注意力实现闭环控制的策略镜像下降

PolicyAttention: Softmax Attention Implements Policy Mirror Descent for Closed-Loop Control

Yuhe Sui, Yingzhi Tang, Shufang Chen

首次发表
浏览论文内容

中文总结 AI 辅助

本研究证明因果软最大注意力可实现策略镜像下降作为闭环控制器,构建固定协议并实验验证,在重复控制任务中损失接近精确PMD,优于其他适应方案。

中文摘要 AI 辅助

因果软最大注意力能否实现策略镜像下降(PMD),作为重复控制器而非一步代数恒等式?负熵策略镜像下降具有逐状态更新 $\operatorname{PMD}_\eta(\pi,Q)=\operatorname{softmax}(\log\pi+\eta Q)$。基于已知的Q-TD-PMD递归,我们构建了一个固定的因果软最大actor-环境-一步critic协议,包含显式的actor、路由、采样和归一化残差,并将它们传播到实际返回的策略。该构造规定了有限logit/全支撑域、外部token化和采样边界,以及归一化编译所需的均值零LayerNorm载体条件。单独训练的pre-LN Transformer在经验上恢复了目标计算。在测试的固定规则中,冻结的一步审计模型最接近PMD;在预注册的五次运行$S=4$重复控制测试中,具有精确一步critic的学习actor达到中位返回策略损失为精确PMD或acles的$1.052$倍,并在四次无重训练偏移中保持该标准。相同的检查点与其学习的critic给出描述性中位$1.050$倍于oracle(无注册边际)。在$S=8$时,将精确critic替换为学习critic将中位$T=20$损失提高到$0.0225$,但使Liang-Lai和算法蒸馏适应方案的损失高出$20.2$-$24.2$倍;这是单边采样critic界限,因为PolicyAttention每轮消耗144个生成性转换,而适应方案仅消耗20个在策略转换。严格的20转换比较仍然开放。在$S=8,16$时,精确critic公共基准比较仍比那些适应方案低$17.7$-$28.2$倍的损失,信息不对称在局部说明。

英文摘要

Can causal softmax attention implement policy mirror descent as a repeated controller rather than a one-step algebraic identity? Negative-entropy policy mirror descent (PMD) has the statewise update $\operatorname{PMD}_η(π,Q)=\operatorname{softmax}(\logπ+ηQ)$. Building on the known Q-TD-PMD recursion, we construct one fixed causal-softmax actor--environment--one-step-critic protocol with explicit actor, routing, sampling, and normalization residuals, and propagate them to the policy actually returned. The construction states the finite-logit/full-support domain, the external tokenization and sampling boundary, and the mean-zero LayerNorm carrier conditions required by the normalized compilation. Separately trained pre-LN Transformers recover the target computation empirically. A frozen one-step audit model is closest to PMD among the tested fixed rules; in a preregistered five-run $S=4$ repeated-control test, the learned actor with an exact one-step critic reaches median returned-policy loss $1.052\times$ the Exact PMD oracle and retains the criterion across four no-retraining shifts. The same checkpoints with their learned critic give descriptive median $1.050\times$ the oracle (no registered margin). At $S=8$, replacing the exact critic by the learned critic raises median $T=20$ loss to $0.0225$ yet leaves the Liang--Lai and Algorithm Distillation adaptations $20.2$--$24.2\times$ higher-loss; this is a one-sided sampled-critic bound because PolicyAttention consumes 144 generative transitions per round versus 20 on-policy transitions for the adaptations. The strict 20-transition comparison remains open. At $S=8,16$, the exact-critic common-harness comparison remains $17.7$--$28.2\times$ lower-loss than those adaptations, with the information asymmetry stated locally.

发表机构

  • Quantitative Research Society(量化研究学会)
  • Nanyang Technological University(南洋理工大学)
  • The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
  • The University of Hong Kong(香港大学)

机构由 AI 辅助整理,请以论文原文为准。

↑