智能体策略-价值审计:在金融LLM智能体中分离转换构成与事件选择
Agent Policy-Value Audit: Separating Transition Composition from Event Selection in Financial LLM Agents
- Emory University(埃默里大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出智能体策略-价值审计方法,通过固定动作转换计数并随机重分配事件,分离金融LLM智能体的部署价值与事件选择价值,纠正传统配对检验的误判。
AI中文摘要:
金融LLM智能体通常通过将其端到端收益与基线进行比较,并对配对差异进行零假设检验来评估。这衡量了部署智能体是否改变了已实现的绩效,但并未隔离事件选择技能。一个频繁将仓位从空头转为多头的智能体,即使随机选择事件,也能从向上漂移的事件池中获得正的配对收益。我们提出了智能体策略-价值审计,该方法固定每种有序动作转换类型的观测计数,并在符合条件的事件中随机重新分配它们。这些重新分配的平均收益是构成基准;观测部署价值与该基准之间的差异是选择价值。在基于真实盈利事件收益的半合成基准中,零中心配对检验在11.6%的无技能重复中错误地将被动暴露归因于选择技能,而转换匹配审计将该比率降至5.3%。对44家美国面向消费者公司的723个盈利事件进行回顾性应用,该审计将智能体+15.2个基点/事件的毛部署价值分解为+25.8个基点的构成基准和-10.6个基点的选择价值。该智能体并未显著优于匹配的随机分配。金融智能体评估应分别报告部署价值和事件选择价值。
英文摘要:
Financial LLM agents are often evaluated by comparing their end-to-end returns with those of a baseline and testing the paired difference against zero. This measures whether deploying the agent changes realized performance, but it does not isolate event-selection skill. An agent that frequently changes positions from flat to long can earn a positive paired return from an upward-drifting event pool even when it selects events at random. We propose the Agent Policy-Value Audit, which holds fixed the observed count of each ordered action-change type and randomly reassigns them across eligible events. The average payoff from these reassignments is the composition benchmark; the difference between observed deployment value and this benchmark is selection value. In semi-synthetic benchmarks based on real earnings-event returns, a zero-centered paired test falsely attributes passive exposure to selection skill in $11.6\%$ of no-skill replications, while the transition-matched audit reduces this rate to $5.3\%$. Applied retrospectively to 723 earnings events at 44 U.S. consumer-facing firms, the audit decomposes the agent's gross deployment value of $+15.2$ bps/event into a $+25.8$ composition benchmark and a $-10.6$ selection value. The agent does not detectably outperform matched random assignments. Financial-agent evaluations should report deployment value separately from event-selection value.