评估大语言模型的投资逻辑:面向个性化金融智能体的真实世界基准
Evaluating Investment Logic in Large Language Models: A Real-World Benchmark Towards Personalzied Financial Agents
浏览论文内容
中文总结 AI 辅助
该研究推出原生过程基准InvestLogicBench,含151位真实投资者的201247项决策,评估发现金融LLMs逻辑合理性高但事件接地性低,指出需以P→E→R→D→O轨迹评估个性化决策智能体。
中文摘要 AI 辅助
投资能力本质上是个性化的:相同的市场证据对具有不同目标、时间范围、投资组合和风险边界的投资者而言,可能会证明不同的行动是合理的。然而,金融大语言模型(LLMs)的评估要么通过静态问答进行,要么通过最终盈亏进行。前者忽略了智能体的属性;后者无法说明盈利行动是否基于投资者的特征、是否符合投资者的情况,或者仅仅是运气使然。我们质疑学术界是否在用错误的标准评估具有重要决策影响的智能体。我们推出\textsc{InvestLogicBench},这是一个原生过程的基准,包含来自151位真实世界投资者的201247个已记录决策。每个案例都实例化一个\textbf{P$\rightarrow$E$\rightarrow$R$\rightarrow$D$\rightarrow$O}轨迹:投资者\textit{Profile}(特征)、可观察的市场\textit{Events}(事件)、投资\textit{Reasoning}(推理)、可执行的\textit{Decision}(决策)和延迟的\textit{Outcome}(结果)。该基准的发布内容包括特征构建、时点事件绑定、结构化逻辑、时间范围、结果和事后分析,支持理解、基于特征条件的生成以及端到端重放。在四个领先的LLMs上,逻辑合理性仍接近4/5,而事件接地性仅为0.8--2.8/5;回报与过程质量也不一致。这些结果揭示了仅基于结果的评估所隐藏的、看似完善但接地性较弱的推理。我们进一步认为,P$\rightarrow$E$\rightarrow$R$\rightarrow$D$\rightarrow$O应成为数据系统的接口,需要带版本的特征、时间溯源、可检查的检索、决策账本和可重放的结果。金融是我们对更广泛的个性化、具有重要决策影响的智能体的压力测试。
英文摘要
Investment competence is inherently personalized: the same market evidence can justify different actions for investors with different goals, horizons, portfolios, and risk boundaries. Yet financial LLMs are evaluated either by static question answering or by terminal profit and loss. The former omits agency; the latter cannot reveal whether a profitable action was grounded, profile-consistent, or merely lucky. We ask whether the community is using the wrong ruler for consequential agents. We introduce \textsc{InvestLogicBench}, a process-native benchmark containing 201,247 documented decisions from 151 real-world investors. Each episode instantiates a \textbf{P$\rightarrow$E$\rightarrow$R$\rightarrow$D$\rightarrow$O} trace: investor \textit{Profile}, observable market \textit{Events}, investment \textit{Reasoning}, executable \textit{Decision}, and delayed \textit{Outcome}. The release includes profile construction, point-in-time event binding, structured logic, horizons, outcomes, and post-mortems, and supports comprehension, profile-conditioned generation, and end-to-end replay. Across four leading LLMs, logical plausibility remains near 4/5 while event grounding is only 0.8--2.8/5; return and process quality also disagree. These results expose polished but weakly grounded reasoning that outcome-only evaluation hides. We further argue that P$\rightarrow$E$\rightarrow$R$\rightarrow$D$\rightarrow$O should be a data-system interface, requiring versioned profiles, temporal provenance, inspectable retrieval, decision ledgers, and replayable outcomes. Finance is our stress test for a broader class of personalized, consequential agents.