arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.29701cs.MA

多智能体辩论用于可解释交易:模拟市场中的推理、共识与表现

Multi-Agent Debate for Explainable Trading: Reasoning, Consensus, and Performance in Simulated Markets

Juli Huang, Alanood Alrassan, Deveen Harischandra, Theodore Wu, Veljko Skarich, Matthew Hayes

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过多智能体辩论框架模拟历史市场投资,发现推理质量提升不直接转化为财务收益,而保留分歧的干预能显著改善夏普比率,强调辩论应保留独立信号。

中文摘要 AI 辅助

大型语言模型(LLMs)越来越多地被用于金融决策,但推理质量的提升是否能转化为更好的经济结果仍不清楚。我们通过一个用于历史市场模拟中投资组合分配的多智能体辩论框架来研究这一问题,其中专门化的智能体提出、批评并修订投资决策。推理质量从四个维度进行评估:逻辑有效性、证据支持、替代方案考虑和因果一致性,并与下游财务表现进行比较。在210次受控运行中,总体推理质量与夏普比率(r = 0.07,p = 0.29)或总回报(r = 0.03,p = 0.70)之间没有显著关系。结构化提示将测得的推理质量从约0.72提高到0.84(+17.7%,Cohen's d约2.0),但这些提升并未持续转化为更高的回报。我们识别出奉承性趋同作为一个核心失败模式,其中智能体在批评-修订循环中放弃独立立场并趋同于相似的分配。一种保持分歧的Jensen-Shannon散度干预将夏普比率提高了+0.14(p = 0.028),Sortino比率提高了+0.25(p = 0.026),而强制更强因果推理的干预并未改善财务表现。我们的结果表明,多智能体辩论在保留独立信息信号时最有价值,而不仅仅是提高测得的推理质量。

英文摘要

Large language models (LLMs) are increasingly used for financial decision-making, yet it remains unclear whether improvements in reasoning quality translate into better economic outcomes. We investigate this question using a multi-agent debate framework for portfolio allocation in historical market simulations, where specialized agents propose, critique, and revise investment decisions. Reasoning quality is evaluated across four dimensions: logical validity, evidential support, alternative consideration, and causal alignment, and compared with downstream financial performance. Across 210 controlled runs, aggregate reasoning quality shows no meaningful relationship with Sharpe ratio (r = 0.07, p = 0.29) or total return (r = 0.03, p = 0.70). Structured prompting increases measured reasoning quality from about 0.72 to 0.84 (+17.7%, Cohen's d about 2.0), but these gains do not consistently translate into higher returns. We identify sycophantic convergence as a central failure mode, where agents abandon independent positions during critique-revision cycles and converge toward similar allocations. A Jensen-Shannon divergence intervention that preserves disagreement improves Sharpe by +0.14 (p = 0.028) and Sortino by +0.25 (p = 0.026), while interventions enforcing stronger causal reasoning do not improve financial performance. Our results suggest that multi-agent debate is most valuable when it preserves independent informational signals rather than simply improving measured reasoning quality.

发表机构

  • Stanford University(斯坦福大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑