发表机构
Karlsruhe Institute of Technology; Leibniz Institute for Primate Research; salestech Data & AI(卡尔斯鲁厄理工学院; 莱布尼茨灵长类动物研究所; salestech数据与人工智能公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过记录四个开放权重模型在144个博弈中的激活,揭示了激励到选择的内部路径因模型而异,表明相似行为可基于不同计算,后训练重塑决策路径而不改变行为。
AI 中文摘要
大型语言模型扮演着战略智能体和人类选择模型的角色,然而,像战略智能体一样选择并不意味着像战略智能体一样计算。我们记录了四个开放权重模型——包括密集型和混合专家模型,以及一对匹配的基础-指令微调模型——在144个严格序数$2\ imes2$博弈的一次性博弈中的激活状态。我们遵循了一个预先指定的激励路径,从提示开始,经过激活,最终到达选择。密集模型反映了人类随博弈复杂性增加而未经调整的衰退趋势。激励和选择在每个模型中都是可检测的,但模型之间的差异在于激励是否到达选择、是否与选择对齐,以及在测试的情况下,加强激励是否改变了偏好。基础版和指令微调版的Qwen2.5模型在基线时选择几乎相同,但在激励是否到达选择方面存在差异。固定的决策线索在内部是可区分的,但选择性地改变了选择。相似的行为可能基于不同的计算;训练后的微调可以重塑从表征激励到决策的路径,同时使行为和可解码信息基本保持不变。
英文摘要
Large language models act as strategic agents and models of human choice, yet choosing like a strategic agent does not mean computing like one. We recorded activations from four open-weight models --- dense and mixture-of-experts, including a matched base--instruct pair --- in one-shot play of 144 strict ordinal $2\times2$ games. We followed a prespecified incentive from prompt, through activations, to choice. Dense models mirrored the unadjusted human decline with game complexity. Incentive and choice were detectable in every model, but models differed in whether incentive reached the choice, aligned with it and, where tested, whether strengthening it shifted preference. The base and instruction-tuned Qwen2.5 models chose almost identically at baseline yet differed in whether incentive reached choice. Fixed decision cues were distinguishable internally but changed choices selectively. Similar behaviour can rest on different computation; post-training can reshape the path from represented incentive to decision while leaving behaviour and decodable information largely intact.
CommentsV2 adds the link for the reproduction package (GitHub)