在维基百科摘要上玩 log(N) 个问题:配对前沿模型之间的通信效率
Playing log(N)-Questions over Wikipedia Abstracts: How Per-Round Errors Compound Under Information Asymmetry
浏览论文内容
中文总结 AI 辅助
本研究通过双智能体 log(N)-Questions 游戏评估六种前沿语言模型的自我通信效率,发现模型间性能差异显著,胜率随集合大小呈幂律下降,并识别出标题划分策略与信息提取的相关性。
中文摘要 AI 辅助
我们在双智能体 log(N)-Questions 游戏中评估了六种前沿语言模型。提问者看到 N 个维基百科首段,必须使用恰好 log2 N 个是/否问题识别出秘密选择的目标。回答者只看到目标和问题,并回复一个词。两个角色在同一提供商上运行,因此该游戏衡量了模型在信息不对称情况下与自身通信的能力。我们在包含 4 到 1024 个段落的文档集上运行了 408 场游戏,总 API 成本为 363 美元。一个模型远远落后于其他模型:Claude Opus 5 在 68 场游戏中赢得 28 场,而 GLM-5.3、GPT-5.6 Sol、Grok 4.6、Gemini 3.8 Flash 和 Kimi K3 则赢得 45 到 56 场。领先的五个模型仅勉强可区分。将这五个模型合并后,胜率随集合大小下降,r=-0.973,并可由单一每轮可靠性参数拟合。其形式为 win=p^(log2 N),其中 p=0.928。失败大致平均分为回答错误和区分失败,模型几乎从不指出其自身证据排除的文档。最弱模型的每个一致回答错误都经过检查:34 个中有 32 个是“否”回答,针对文档第一句中陈述的属性,而指令明确警告不要默认回答“否”。根据回答平衡估计的每问题信息量与胜率相关,r=+0.88。仅有的两个每问题提取完整比特的模型是仅有的两个按文档标题划分的模型,这种策略在 N<32 时不存在,在 N≥32 时用于四分之一的问题。推理 token 消耗在模型间变化 4.5 倍,与成功关系不大,并且随着候选集缩小,轨迹增长而可靠性没有相应提高。
英文摘要
We evaluate six frontier language models on the two-agent $\log_2 N$-Questions game (Potash et al., 2019) to measure self-communication across an information asymmetry. A questioner with access to $N$ candidate Wikipedia lead paragraphs ($N = 4$ to $1024$) must identify a secret target using exactly $\log_2 N$ binary questions answered by an agent from the same provider that sees only the target. Across 408 games, win rate decays cleanly as a geometric power of horizon length, $p^{\log_2 N}$ ($p \approx 0.93$). Per-round failure rates are flat across the horizon, indicating that errors compound because more rounds must succeed rather than because individual rounds grow harder. Adjudication across three independent judges shows that losses divide between single-agent answer errors and discrimination failures, which become undetectable and unrecoverable under the two-agent structure rather than from channel breakdown. Claude Opus 5 lags behind due to systematic false-negative answers (82% answer errors), whereas the five leading models (GLM-5.3, GPT-5.6 Sol, Grok 4.6, Gemini 3.8 Flash, and Kimi K3) are closely clustered. Maximizing information gain requires structural partitioning (e.g., splitting on document titles), and neither reasoning-token expenditure nor API cost correlates with success ($r = -0.05$), highlighting communicative reliability as a distinct bottleneck from inference compute.