arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.04373cs.AIcs.CY

为何更优模型会创造更具风险的系统:来自金融市场中大型语言模型(LLM)智能体的证据

Why Better Models Can Create Riskier Systems: Evidence from LLM Agents in Financial Markets

  • Massachusetts Institute of Technology(麻省理工学院)
  • Harvard Kennedy School(哈佛肯尼迪学院)

机构由 AI 辅助整理,请以论文原文为准。

Jillian Ross, Eric So, Zoe De Simone, Charles Pozniak, Andrew W. Lo

AI总结:

该研究以金融市场LLM智能体为对象,发现提升单个LLM能力可能增强其行为关联性,在特定情况下会增加系统风险,揭示了能力悖论,为相关系统设计提供了新视角。

AI中文摘要:

大型语言模型(LLM)正被大规模部署于金融市场、内容审核、招聘等关键现实系统中。本文表明,提升单个模型的能力可能会降低而非改善系统层面的结果。我们假设,共享训练与架构会使更强大的LLMs表现出更相似的行为,产生无法通过多样化消除的关联行动。我们开发了一个通用框架,展示这种关联如何形成不可多样化的风险底线,并使用具有不同通用能力的LLM交易者的基于智能体的模拟,在金融市场中验证其预测。我们发现:(1)前沿LLMs表现出显著的关联行为,且该行为随能力提升而增强;(2)当它们的共享推理准确时,增加智能体参与度会降低市场层面的风险;(3)当智能体处于共同的错误信息环境中时,相同的关联行为会成为负担。这些结果共同揭示了一个能力悖论:提升单个模型的能力不一定会产生更好的系统层面结果。这种动态是否会在其他领域出现是一个开放的实证问题。

英文摘要:

Large language models (LLMs) are being deployed at scale in consequential real-world systems, from financial markets to content moderation to hiring. We show that improving individual model capability can degrade rather than improve system-level outcomes. We hypothesize that shared training and architectures can lead more capable LLMs to behave more similarly, creating correlated actions that do not diversify away. We develop a general framework showing how this correlation creates a non-diversifiable risk floor and test its predictions in financial markets using an agent-based simulation with LLM traders of varying general-purpose capability. We find that: (1) frontier LLMs exhibit significantly correlated behavior that increases with capability; (2) when their shared reasoning is accurate, increasing agent participation reduces market-level risk; and (3) when agents share a common misinformation environment, the same correlated behavior becomes a liability. Together, these results identify a capability paradox: improving individual models does not necessarily produce better system-level outcomes. Whether the same dynamics arise in other domains is an open empirical question.

补充信息

↑