arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.09988cs.CEcs.CL

OpenPM:面向LLM投资组合管理智能体的可审计时点评估框架

OpenPM: Auditable Point-in-Time Evaluation for LLM Portfolio-Management Agents

Xinying Cai, Minghao Guo, Jiahe Liu, Jiaojiao Han, Bangwei Guo, Yitao Long, Yuxuan Chen, Bohan Wu, Dimitris N. Metaxas, Raymond Li

首次发表
浏览论文内容

中文总结 AI 辅助

OpenPM是面向LLM投资组合管理智能体的可审计时点评估框架,构建了分层分配器参考智能体,发现分析师质量比构造器选择更重要,换手率是主要成本驱动因素。

中文摘要 AI 辅助

大型语言模型正越来越多地被用于解读市场、评估风险和分配资本。然而,LLM交易智能体的报告结果可能因前瞻泄露、乐观执行以及仅被描述但未被执行的风险指令而被夸大。我们提出了OpenPM,这是一个面向LLM投资组合管理智能体的可审计时点评估框架。在OpenPM中,智能体在标普500(S&P 500)范围内管理100万美元的多头投资组合,使用五分钟间隔的市场数据。智能体可见的每条记录都必须在决策时可用。自然语言风险指令被转换为类型化约束,并在执行的投资组合上强制执行。每次运行都会生成审计工件,包括污染证书、成本敏感性曲线和约束遵守报告。我们还构建了一个名为分层分配器(tiered allocator)的参考智能体,其中类型化分析师对候选资产进行评分,一个构造器LLM提出权重,一个确定性评论家保证可行性。我们通过捕获分析师证据一次并在多个构造器模型中重放,来隔离构造器的行为。在我们的短窗口案例研究中,更强的构造器在同一资产池上相对于等权重表现出适度且依赖模型的收益,但分析师质量比构造器选择更重要,换手率是主要的成本驱动因素。所有收益都是单个冻结窗口的上限,不考虑市场冲击,也未经过验证的阿尔法(alpha)。

英文摘要

Large language models are increasingly used to read markets, assess risk, and allocate capital. However, reported results for LLM trading agents can be inflated by look-ahead leakage, optimistic execution, and risk mandates that are described but not enforced. We present OpenPM, an auditable point-in-time evaluation framework for LLM portfolio-management agents. In OpenPM, an agent manages a \$1M long-only book over the S\&P 500 universe using market data at five-minute intervals. Every record visible to the agent must be available at the decision time. Natural-language risk mandates are converted into typed constraints and enforced on the executed portfolio. Each run produces audit artifacts, including a contamination certificate, a cost-sensitivity curve, and a constraint-adherence report. We also build a reference agent named the tiered allocator, where typed analysts score candidates, a constructor LLM proposes weights, and a deterministic critic guarantees feasibility. We isolate constructor behavior by capturing analyst evidence once and replaying it across constructor models. In our short-window case study, stronger constructors show modest and model-dependent gains over equal weighting on the same pool, but analyst quality matters more than constructor choice, and turnover is the main cost driver. All returns are upper bounds on a single frozen window without market impact, not validated alpha.

发表机构

  • Rutgers University(罗格斯大学)
  • Technical University of Denmark(丹麦技术大学)
  • New Jersey Institute of Technology(新泽西理工学院)
  • New York University(纽约大学)
  • Columbia University(哥伦比亚大学)
  • University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
  • University of British Columbia(不列颠哥伦比亚大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑