arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

约定差距:迈向合作AI评估中的隐式沟通度量

The Convention Gap: Towards Measuring Implicit Communication in Cooperative AI Evaluation

Makoto Fukushima, Hua-Dong Xiong, Ehsan Moradi Pari

arXiv 2609.11489首次发表:更新:

发表机构

Honda Research Institute Japan; Honda Research Institute USA(本田日本研究院; 本田美国研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出“约定差距”指标,通过对比字面预测与实际失败率,衡量合作AI中的隐式沟通,并在Hanabi游戏中验证,发现人类与AI的约定兼容性比AI-AI性能更能预测合作效果。

AI 中文摘要

合作型AI智能体通常与其他AI进行对抗评估,然而人类的合作依赖于隐式约定——即超越字面信息解读含义的共享协议——而AI-AI基准测试可能无法捕捉到这一点。我们提出了“约定差距”这一度量,即根据沟通的字面内容预测的失败概率与观察到的实际失败率之间的差异,用以衡量隐式沟通。在纸牌游戏Hanabi中,有限的牌堆和确定性的提示约束使得这一后验概率可以精确计算。我们重放了来自三个人类-人类(此http URL)、AI-AI(HOAD)和人类-AI(HanabiData)公开数据集的约101,000个游戏动作。结果显示,人类配对中的差距为+26.2个百分点(pp),AI配对中为-0.7 pp,人类-AI配对中为+16.4 pp,且差距集中在未收到提示的牌上(人类配对中为+46 pp)。在人类-AI游戏中,人类可获得的字面信息在三个AI伙伴之间相似(平均预测失败率为38%至41%),但人类的实际失败率从14.4%到34.4%不等,差距从+24.1 pp到+6.2 pp;引发最大差距的伙伴导致的人类失败次数最少。游戏得分携带了不同信息:它取决于每个语料库的玩家构成,而差距在智能体层面将人类与AI的游戏行为区分开来。作为已知答案的检验,Off-Belief Learning智能体(其约定内容通过构造得以控制)在无约定水平上给出了+1.6 pp的差距,并单调上升至+21.7 pp。这些结果表明,约定兼容性而非AI-AI性能,可能更能预测AI与人类伙伴合作的有效性。

英文摘要

Cooperative AI agents are evaluated against other AIs, yet human cooperation relies on implicit conventions -- shared protocols for reading meaning beyond the literal message -- which AI-AI benchmarks may not capture. We propose the convention gap, the difference between the failure probability predicted from the literal content of communication and the observed failure rate, as a metric of implicit communication. In the card game Hanabi, the finite deck and deterministic hint constraints make this posterior exactly computable. We replayed about 101,000 play actions from three public datasets of human-human (an online Hanabi platform), AI-AI (HOAD), and human-AI (HanabiData) games. The gap was +26.2 percentage points (pp) in human pairs, -0.7 pp in AI pairs, and +16.4 pp in human-AI pairs, and was concentrated on plays of cards that had received no hints (+46 pp in human pairs). Within human-AI play, the literal information available to humans was similar across the three AI partners (mean predicted failure 38-41%), but human failure rates ranged from 14.4% to 34.4% and the gap from +24.1 to +6.2 pp; the partner eliciting the largest gap produced the fewest human failures. Game score carried different information: it depended on each corpus's roster composition, whereas the gap separated human from AI play at the agent level. As a known-answer check, Off-Belief Learning agents, whose convention content is controlled by construction, gave a gap of +1.6 pp at the convention-free level, rising monotonically to +21.7 pp. These results suggest that convention compatibility, rather than AI-AI performance, may predict an AI's effectiveness with human partners.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑