提议而非评判:面向挖掘投资因子的大语言模型智能体的任意时点有效裁判
Propose, Don't Judge: An Anytime-Valid Referee for LLM Agents That Mine Investment Factors
查看机构详情
- DeepGrounding
- AlphaAvatar
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本研究提出由不可触及的冻结统计裁判对智能体提出的投资因子进行投注式评判,确保任意时点有效控制错误发现率,实验表明冻结裁判显著减少虚假接纳,且语言模型在提议和探针编写上表现最佳。
中文摘要 AI 辅助
语言模型智能体现已承担量化因子研究的全部流程:它们提出投资因子、进行回测、筛选幸存者并淘汰失效因子。我们探讨智能体应保留其中哪些职责。我们的答案是受控的自我进化:智能体可以提议,而一个智能体无法触及的冻结统计裁判必须进行评判。该裁判仅根据提交后显现的市场结果,通过投注方式对每个候选因子评分,因此其错误发现保证在任意停止时间对任何提议策略均成立。我们将三种提议者(脚本、bandit算法和语言模型)与该裁判以及三种故意泄漏的裁判进行交叉对比,实验环境包括一个具有植入真相的合成世界、一个探针编写环境以及基于中证500的十年向前推进测试。评判者决定了虚假接纳的数量:在脚本提议者下,冻结裁判接纳的低于阈值因子数量比泄漏裁判少5至11倍,且没有任何提议者能弥合这一差距。提议者决定了产出:语言模型优于脚本,与bandit算法持平,并增加了bandit所缺乏的一项能力——编写其自身的诊断探针。该认证的代价是时间:一个被接纳的真实因子需等待约500个交易日,因此认证组合的夏普比率落后于未认证组合。评判属于程序;提议和工具制造属于智能体。
英文摘要
Language-model agents now run the whole of quantitative factor research: they propose investment factors, backtest them, select the survivors and retire them. We ask which of those jobs an agent should keep. Our answer is governed self-evolution: the agent may propose, and a frozen statistical referee that the agent cannot touch must judge. The referee scores each candidate only on market outcomes revealed after submission, by betting, so its false-discovery guarantee holds at every stopping time for any proposal policy. We cross three proposers (a script, a bandit and a language model) with this referee and with three deliberately leaky ones, in a synthetic world with planted truth, a probe-authoring environment and a ten-year walk-forward on the CSI 500. Who judges sets the number of false admissions: the frozen referee admits 5-11 times fewer sub-threshold factors than the leaky referees under a scripted proposer, and no proposer closes that gap. Who proposes sets the yield: the language model beats the script, matches the bandit, and adds the one capability a bandit lacks, writing its own diagnostic probes. The certificate's price is time: an admitted true factor waits about 500 trading days, and the certified portfolio's Sharpe ratio therefore trails an ungated one. Judging belongs to the procedure; proposing and instrument-making belong to the agent.