arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.27734q-fin.ST

经得住诚实评估的是什么?针对大语言模型驱动的交易策略发现的防泄漏、感知搜索的评估方法

What survives honest evaluation? Leakage-safe, search-aware assessment of LLM-driven trading strategy discovery

Eray Gençay

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对LLM发现交易策略时的前瞻偏差与搜索强度未校正问题,提出防泄漏的结构防护及基于试验次数的性能调整系统,经实证验证其可诚实评估策略并量化可信验证所需样本量。

中文摘要 AI 辅助

大语言模型(LLM)越来越多地被用于发现交易策略,相关文献存在一个普遍的方法学缺陷:生成大量候选策略,仅报告最优策略,且未校正前瞻偏差,也未校正报告结果背后搜索过程的强度。我们提出一种策略发现系统,从结构层面而非流程层面进行这两项校正。首先,智能体仅能通过注册验证的工具执行操作,这些工具的特征空间从构造上排除了前瞻偏差;我们表明这一防护措施与统计校正并非冗余:一个刻意泄漏的 oracle 公布的夏普比率达35,在经调整的夏普比率和回测过拟合概率测试中完全通过。其次,系统记录搜索过程中执行的每一次策略评估,并根据该试验次数对所有报告的性能进行调整,追踪样本内最优夏普比率如何随每次试验上升,而由智能体自身搜索驱动的调整阈值上升更快。在包含453只股票的时点美国股票 universe 和包含39只ETF的多资产 universe 中,考虑现实的交易、冲击和借贷成本后,诚实评估验证了被动基准(样本外置信区间不包含零),拒绝了所有LLM发现的策略(涉及两个前沿模型、最高100个候选策略的搜索预算、五次重复运行),通过互补工具捕捉了选择运气、预测排名退化和样本外崩溃,并在相同工具下评估了人类交易员的生产规则系统。该框架明确了为何预注册假设比暴力搜索获得更低的证据门槛,并量化了可信验证适度优势所需的样本量。

英文摘要

Large language models (LLMs) are increasingly used to discover trading strategies, and much of the resulting literature shares a methodological weakness: many candidate strategies are generated, the best is reported, and neither look-ahead bias nor the intensity of the search behind the reported result is corrected for. We present a strategy-discovery system that makes both corrections structural rather than procedural. First, the agent can only act through registry-validated tools whose feature space excludes look-ahead by construction; we show that this guardrail is not redundant with statistical correction: a deliberately leaky oracle posting a Sharpe ratio of 35 survives Deflated Sharpe and probability-of-backtest-overfitting testing completely. Second, the system records every strategy evaluation its search performs and deflates all reported performance by that trial count, tracing how the best in-sample Sharpe ratio climbs with each trial while the deflation threshold, driven by the agent's own search, climbs faster. Across a 453-stock point-in-time US equity universe and a 39-ETF multi-asset universe with realistic transaction, impact, and borrow costs, honest evaluation certifies passive benchmarks (out-of-sample confidence intervals excluding zero), rejects every LLM-discovered strategy (across two frontier models, search budgets up to one hundred candidates, and five repeated runs), catching selection luck, predicted rank degradation, and out-of-sample collapse through complementary instruments, and evaluates a human trader's production rule system under identical instruments. The framework formalizes why pre-registered hypotheses earn lower evidential bars than brute search, and quantifies the sample sizes that credible certification of moderate edges actually requires.

↑