arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Amadeus:当人的模型相遇

Amadeus: When Models of People Meet

Karl Hanna

arXiv 2609.35835首次发表:更新:

发表机构

Queen’s University Belfast(贝尔法斯特女王大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文通过国际象棋对局,独立学习8位精英棋手模型并组合,验证了模型能恢复未见交互的某些方面,且不同行为方面的恢复可兼得。

AI 中文摘要

随着人工智能领域持续不断的飞速发展,一个可能浮现于我们脑海的问题是:能否以人类为蓝本对智能体进行建模,并进而利用这些智能体执行合成交互,以预测其真实对应者的行为,甚至预测更大规模如群体或社会的交互?在本文中,我们通过国际象棋对这一问题的一个更受控版本进行了测试。我们选取了8位精英棋手,封存他们之间的直接两两对局,使用不同方法独立学习每位棋手,然后在被保留的对局组合上组合所得模型。为评估生成的交互,我们采用两种度量:开局家族全变差距离和胜-负-和(WDL)全变差距离。M1方法降低了WDL-TV,而开局家族TV基本不变;M2方法则大幅降低开局家族TV,而对WDL-TV影响甚微。一种结合前两者组件的额外事后方法在两个度量上均保持改进。这些结果表明,独立学习的模型能够恢复先前未见交互的某些方面,且不同行为方面的恢复并非互斥。

英文摘要

With the constant advancements in AI, one possibility is to model agents after humans and, in turn, use these agents to carry out synthetic interactions. Such models could be used to predict interactions between their real counterparts, or potentially interactions at larger scales. In this paper, we test a more controlled version of this question through chess. We use 8 elite chess players, seal their direct pairwise games, learn each player independently using different methods, and then compose the resulting models on the withheld dyads. To evaluate the generated interactions, we use two measurements: opening-family total variation distance and win-draw-loss (WDL) total variation distance. M1 primarily improves WDL fidelity while producing smaller opening-family improvements, whereas M2 produces much larger opening-family improvements while having little effect on WDL-TV. For opening-family behaviour under M2, the correct assignment of the eight learned player identities also gives the closest match among all $8! = 40{,}320$ possible assignments. These results show that at least some properties of previously unseen interactions can be recovered from independently learned individuals. The partial recovery observed here may reflect limitations of the current individual modelling methods rather than a fundamental limit on compositional interaction recovery. An additional post-hoc method that combines the two mechanisms improves both measurements, suggesting that recovery across these behavioural properties is not necessarily mutually exclusive.

CommentsSubmitted to AAMAS 2027

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑