发表机构
Stanford University; Nanjing University; Hasso Plattner Institute(斯坦福大学; 南京大学; 哈索·普拉特纳研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出智能体实证资产定价(AEAP)范式,明确其核心构建模块,构建因子发现的参考架构与评估标准,通过SEADS实验发现单一指标无法稳定评估系统,还揭示了AEAP系统的评估陷阱。
AI 中文摘要
大型语言模型(LLM)智能体的最新进展催生了一种新的资产定价范式,我们将其称为智能体实证资产定价(Agentic Empirical Asset Pricing, AEAP):即能够自主开展科学发现过程的系统。我们定义了AEAP并确定了其核心构建模块。现有评估实践仅对产出(因子或交易)进行回测,而非对产生这些产出的自主发现系统进行评估。我们聚焦于因子发现,贡献了参考架构、针对发现因子的严格评估标准,以及对发现系统进行样本外回测的方法。作为该架构的具体实例,我们使用此标准在两个美国股票面板上针对五个重新实现的基线评估了SEADS:没有任何单一指标能始终如一地对系统进行排名,这促使我们同时从多个维度进行评估。随后的独立滚动重新执行则提出了一个补充性问题:是发现过程本身,而非某个静态产出,是否可靠。我们还报告了负面发现和局限性,这些问题为未来的AEAP系统揭示了进一步的评估陷阱。
英文摘要
Recent advances in LLM agents enable a new paradigm for asset pricing, which we call Agentic Empirical Asset Pricing (AEAP): systems that autonomously conduct the scientific discovery process itself. We define AEAP and identify its core building blocks. Existing evaluation practices backtest only the outputs (factors or trades), not the autonomous discovery system that produced them. We focus on factor discovery, contributing a reference architecture, a rigorous evaluation standard for discovered factors, and a method for out-of-sample backtesting the discovery system. As a concrete instance of that architecture, we evaluate SEADS against five re-implemented baselines on two US equity panels using this standard: no single metric ranks the systems consistently, motivating evaluation on multiple axes at once. A separate rolling re-execution then asks the complementary question of whether the discovery process itself, not one static output, is reliable. We also report negative findings and limitations that surface further evaluation pitfalls for future AEAP systems.
Comments26 pages, 5 figures, 12 tables