arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

几何感知位置编码是否有助于Transformer在空间不完美信息博弈中表现?

Do Geometry-Aware Positional Encodings Help Transformers in Spatial Imperfect-Information Games?

Wenji Fu

arXiv 2608.14982首次发表:更新:

发表机构

Research Institute of Economics and Management; Southwestern University of Finance and Economics(经济与管理研究院; 西南财经大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究在六角形海军追捕游戏基准上验证,HexRoPE位置编码可提升Transformer的信念估计与数据高效模仿能力,但未提高整体游戏胜率。

AI 中文摘要

应用于空间不完美信息博弈的Transformer需要表示地图几何结构,同时随时间跟踪隐藏实体。本文探究几何感知位置编码是否能提升这些能力,且不声称提出新的位置编码。我们在六角形海军追捕游戏上构建了四级基准:受控几何与拓扑探针、精确贝叶斯隐藏目标跟踪任务、1000场和10000场游戏的离线策略模仿,以及针对三个传统对手的7200场固定种子游戏。在匹配的Transformer骨干网络中,HexRoPE在D6变换的测试轨道上相较于无位置编码降低了精确信念后验交叉熵0.278,在更大地图上降低0.329;两者的分层自举置信区间均不包含零,霍尔姆校正后的p值均低于0.001。在1000场游戏中,HexRoPE的策略动作准确率相较于无编码提升4.63个百分点,相较于矩形相对偏差提升2.05个百分点;在10000场游戏中,增益分别缩减至1.55和0.41个百分点。然而,HexRoPE并未提升整体游戏胜率:其相较于无编码的配对效应为-1.56个百分点(95%置信区间[-4.50, 1.17])。矩形相对偏差在D6信念一致性上表现最强,但在从半径3外推至半径4时表现急剧下降,而图偏差仅提供少量的阻塞边增益。结果表明,几何归纳偏差可提升信念估计和数据高效模仿,但这些表示增益不会自动产生更强的闭环游戏表现。

英文摘要

Transformers applied to spatial imperfect-information games must represent map geometry while tracking hidden entities through time. We ask whether geometry-aware positional encodings improve these capabilities, without claiming a new positional encoding. We construct a four-level benchmark on a hexagonal naval pursuit game: controlled geometry and topology probes, an exact-Bayes hidden-target tracking task, offline policy imitation at 1k and 10k games, and 7,200 fixed-seed games against three legacy opponents. Across matched Transformer backbones, HexRoPE reduces exact-belief posterior cross-entropy relative to no positional encoding by 0.278 on D6-transformed test orbits and 0.329 on a larger map; both hierarchical-bootstrap confidence intervals exclude zero, and both Holm-adjusted p-values are below 0.001. At 1k games, HexRoPE improves policy action accuracy by 4.63 percentage points over no encoding and 2.05 points over rectangular relative bias; the gains shrink to 1.55 and 0.41 points at 10k games. However, HexRoPE does not improve aggregate gameplay win rate: its paired effect over no encoding is -1.56 percentage points (95% CI [-4.50, 1.17]). Rectangular relative bias is strongest on D6 belief consistency but fails sharply when extrapolating from radius 3 to radius 4, while graph bias provides only a small blocked-edge gain. The results show that geometric inductive bias improves belief estimation and data-efficient imitation, but those representation gains do not automatically produce stronger closed-loop play.

Comments7 pages, 4 figures, 3 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑