发表机构
Independent researcher(独立研究者)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究证明在知情对手共享信号通道时,稳健最优信号与显著性极点重合,通过引入说服预算β的对手模型,揭示从贝叶斯判别转向显著性的机制,并指出评估的结构性限制。
AI 中文摘要
当知情对手共享受限信号通道的受众时,最能保护真相的信号就是最能描述真相的信号。在108个确认性项目上,对抗稳健最优与先前工作中的显著性极点完全一致。在200,000个项目的池中,两者仅在2,748个项目上存在差异——恰好位于先前的显著性到贝叶斯坐标未定义之处。在定义之处,稳健性通过完全从贝叶斯判别转向显著性来实现。我们通过将对手引入一个强制选择任务(从《欺骗:香港谋杀案》中抽象而来)来展示这一点。对手知道目标,观察信号,并使用说服预算β为最强错误答案进行论证。随着β的增长,最优信号从后验最大化选项转向边际最大化选项;在β=0时,游戏再现了具有听众温度τ=1的原始预言机模型。这种效应是真实的:池中18.2%的项目在有限预算下其最优值发生转移,且每个项目的临界预算都是精确的。这种巧合在结构上限制了实证评估。两种对手框架在108个项目中的30到77个项目上改变了七个语言模型的选择,而精确的无效应率为零。然而,没有任何测量能够确定这种移动是朝向对手感知最优还是朝向显著性,因为这两个选项是相同的。这是一个结构性限制,而非零结果。诊断检查成本低廉:在评估对手感知性之前,先验证稳健目标是否与评估项目上的启发式目标一致。
英文摘要
When an informed adversary shares the audience of a constrained signalling channel, the signal that best protects the truth is the signal that best describes it. On 108 confirmatory items, the adversary-robust optimum aligns exactly with the salience pole from prior work. Across a 200,000-item pool, the two differ on only 2,748 items --- lying exactly where the prior salience-to-Bayes coordinate is undefined. Where defined, robustness is achieved by moving from Bayesian discrimination entirely to salience. We show this by introducing an adversary to a forced-choice task (abstracted from Deception: Murder in Hong Kong). The adversary knows the target, observes the signal, and argues for the strongest wrong answer using a persuasion budget, $β$. As $β$ grows, the optimal signal shifts from the posterior-maximizing option to the margin-maximizing one; at $β= 0$, the game reproduces the original oracle model with a listener temperature of $τ= 1$. This effect is real: 18.2 percent of the pool has an optimum that shifts under a finite budget, and each item's critical budget is exact. This coincidence structurally limits empirical evaluation. Two adversary framings change the chosen option of seven language models on 30 to 77 of 108 items against an exact no-effect rate. Yet, no measurement can determine whether this movement is toward the adversary-aware optimum or toward salience, because the two options are identical. This is a structural limit, not a null result. The diagnostic check is cheap: before evaluating adversary-awareness, verify whether the robust target coincides with a heuristic target on the evaluation items.
Comments11 pages, 3 figures