arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

多智能体语言模型中的道德风险

Moral Hazard in Multi-Agent Language Models

Dane Malenfant

arXiv 2607.23982首次发表:更新:

发表机构

Mila - The Québec AI Institute(米拉-魁北克人工智能研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究多智能体语言模型中的道德风险,引入对话道德风险游戏,评估七个模型并分解行为,用多种优化机制更新,发现效果异质,强调应报告机制级行为而非仅团队成功。

AI 中文摘要

当具有社会价值的努力成本高昂、难以观察且主要惠及他人时,合作可能会失败。借鉴霍尔姆斯特伦的团队道德风险模型,我们引入了对话道德风险游戏,这是一种可控的文本游戏,将这种隐藏行动结构应用于语言智能体。在每一轮中,智能体可以保留即时的局部奖励,或支付查询成本以揭示主要帮助另一个智能体下游决策的隐藏安全事实。我们评估了七个开放权重语言模型,并将行为分解为查询使用、实际信息传递、局部奖励保留、不安全选择、格式有效性和团队成功。基础模型通常保留局部奖励但没有团队成功,或进行查询但不传递改变最终决策的信息。然后我们使用监督微调、RLOO、顺序SFT+RLOO和GEPA提示优化作为诊断更新机制。它们的效果是异质的:OLMo-7B显示出最明显的与机制一致的权重级改进,而GEPA有时在提高团队成功的同时减少或消除了成本高昂的查询。因此,优化可以在不恢复预期合作机制的情况下转移总奖励,这促使评估报告机制级行为而非仅团队成功。

英文摘要

Cooperation can fail when socially valuable effort is costly, hard to observe, and benefits mainly someone else. Building on Holmstrom's model of moral hazard in teams, the Dialogue Moral Hazard Game instantiates this hidden-action structure as a textual environment for language agents. An agent chooses between keeping an immediate local reward and paying a query cost to reveal a hidden safety fact that helps another agent's downstream decision. We evaluate fourteen open-weight and four frontier models using measures of information acquisition, communication, downstream use, and team success. In matched 3,015-decision-per-model experiments, GPT-5.6 Sol, Claude Opus 4.8, and Nemotron-3 Ultra track the derived private-share boundary across nine query costs, with mean absolute errors of 0.013, 0.030, and 0.024. Muse Spark 1.1 responds directionally, whereas Fable 5 remains query-saturated. SFT, RLOO, SFT+RLOO, and GEPA produce heterogeneous mechanism changes. GEPA raises Muse team success from 22.2% to 100.0% while reducing query use from 51.1% to 0.3%. Frozen-prompt interventions show that this success depends on a learned rank-label mapping rather than direct revelation: changing the mapping reduces team success from 100.0% to 12.5% and then 0.0%. We introduce CREDIT (Counterfactual Replay for Evidence-Driven Information Transfer), a mechanism-aligned multi-agent prompt-optimization algorithm that uses matched hidden-state twins and total-action replay to reward robust causal contribution rather than query frequency. Across five models and multiple seeds, CREDIT preserves query-mediated behavior while revealing model-specific acquisition and downstream-use bottlenecks. Optimization can reach the same aggregate outcome through direct revelation or a learned effective information structure, motivating mechanism-level evaluation and optimization rather than team success alone.

CommentsProvenance experiments in the appendix, Qwen 27B SFT results, Olmo 32B SFT and SFT + RLOO results

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑