AI 中文总结
研究无范围限注德州扑克自回归模型中隐藏状态探测的模糊性,发现对手范围探测在部分种子中呈阳性,行为头预测有提升。可见投注组合比残余隐藏状态解释更多对手范围信号,提出组合受限的预测支持,强调信念探测需经替代方案解释。
AI 中文摘要
隐藏状态探测通常能在不完全信息序列模型中恢复潜在标签,但这本身并不能确定模型是否在隐藏状态上维持后验信念分布。本文研究了仅在行动和价值目标上训练,而非对手手牌或范围的无范围限注德州扑克自回归模型中的这种模糊性。对手范围探测在三个种子中的两个中,在行动/价值控制后呈阳性,行为头仅使用可观察的公共历史预测未观察到的行动比基线高约五个百分点。然而,可见的公共投注组合比残余隐藏状态解释了更多对手范围信号,表明大多数可恢复信息来自投注总结。行动/价值+组合基线达到16.5-16.7%的前10准确率,而组合-残余隐藏探测降至11.4-12.2%,且每个种子中的匹配组合比较均为阴性。我们称这种证据模式为组合受限的预测支持:隐藏状态保持行为预测性和对手范围相关性,但大多数可恢复范围信息由可见投注组合而非残余隐藏状态结构解释。这是关于对手范围表征证据的案例研究主张,而非精确的贝叶斯后验跟踪或因果信念机制。综合控制和预言验证表明,在匹配控制下,相同诊断接受后验敏感状态并拒绝原始组合状态。因此,在将积极信念探测视为信念跟踪证据之前,应通过有针对性的替代方案进行解释。
英文摘要
Hidden-state probes often recover latent labels in imperfect-information sequence models, but this alone does not establish that a model maintains a posterior belief distribution over hidden states. This paper studies this ambiguity in a no-range Limit Hold'em autoregressive model trained only on action and value targets, not on an opponent's hand or range. Opponent-range probes are positive after action/value controls in two of three seeds, and the behavior head predicts held-out actions about five percentage points above a baseline using only observable public history. However, visible public betting composition explains more opponent-range signal than residual hidden states, suggesting that most recoverable information comes from betting summaries. Action/value+composition baselines reach 16.5-16.7% top-10 accuracy while composition-residual hidden probes fall to 11.4-12.2%, and matched-composition comparisons are negative in every seed. We call this evidence pattern composition-bounded predictive support: hidden states remain behavior-predictive and opponent-range correlated, but most recoverable range information is explained by visible betting composition rather than residual hidden-state structure. This is a case-study claim about opponent-range representational evidence, not exact Bayesian posterior tracking or a causal belief mechanism. Synthetic control and oracle validations show that the same diagnostics accept posterior-sensitive states and reject raw composition states under matched controls. Thus positive belief probes should be interpreted through targeted alternatives before being treated as evidence of belief tracking.