arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

在单层自注意力模型中使用(交换)遗憾损失进行训练:概率单纯形的案例研究

Training with (Swap) Regret Loss in a Single-Layer Self-Attention Model: A Case Study on the Probability Simplex

Chanwoo Park, Asuman Ozdaglar

arXiv 2607.23333首次发表:更新:

发表机构

MIT EECS(麻省理工学院电子工程与计算机科学系)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究通过概率单纯形策略视角重审视遗憾损失框架,用其训练单层自注意力模型,新引入交换遗憾损失函数,表明基于遗憾训练的注意力能实现可微机制,使最小注意力架构具备博弈论保证的在线学习动态。

AI 中文摘要

我们通过概率单纯形策略的视角,重新审视了Park等人(2025年)引入的遗憾损失框架,该框架使用决策理论遗憾作为直接损失函数来训练模型以做出更好的决策。我们的第一个结果表明,用遗憾损失训练的单层自注意力模型存在一个驻点,其前向传递与具有适当步长的平滑虚拟博弈完全匹配,确保无悔行为。同时,我们还新引入了交换遗憾损失函数,将遗憾损失框架扩展到外部遗憾之外,使模型能够直接针对交换偏差鲁棒性进行优化。我们进一步表明,这种交换遗憾损失存在一个驻点,其前向传递实现了由经典Blum-Mansour无传递实现算法诱导的相应交换遗憾更新。这些结果表明,基于遗憾训练的注意力可以实现可微机制,其部署在博弈中诱导均衡行为。

英文摘要

We revisit the regret loss framework introduced in Park et al. (2025), which uses decision-theoretic regret as a direct loss function for training models to make better decisions, through the lens of probability-simplex policies. Our first result shows that a single-layer self-attention model trained with regret loss admits a stationary point whose forward-pass exactly matches smoothed fictitious play with the appropriate stepsize that ensures no-regret behavior-i.e., for any given policy input, the model outputs the same update that smoothed fictitious play would produce. In parallel, we also newly introduce a swap-regret loss function, which extends the regret-loss framework beyond external regret and enables models to directly optimize for swap-deviation robustness. We further show that this swap-regret loss admits a stationary point whose forward pass implements the corresponding swap-regret update induced by classical Blum-Mansour no-pass implementation algorithm, with each head implementing an external-regret update via smoothed fictitious play. Together, these results show that regret-trained attention can realize differentiable mechanisms whose deployment induces equilibrium behavior in games: external-regret dynamics lead to coarse correlated equilibrium, while swap-regret dynamics lead to correlated equilibrium. Thus, regret-based objectives steer minimal attention architectures toward online-learning dynamics with game-theoretic guarantees, without supervised traces of those algorithms.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑