arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.03080cs.SEcs.CLcs.LGq-fin.TR

MintEval:大语言模型是否实现了您要求的交易策略?一个面向自然语言到策略代码的行为等价基准

MintEval: Do LLMs Implement the Trading Strategy You Asked For? A Behavioural-Equivalence Benchmark for Natural-Language-to-Strategy Code

Siyu Wang, Yifan Wang, Yuecheng He

首次发表
浏览论文内容

中文总结 AI 辅助

针对大语言模型生成的交易策略代码可能静默偏离用户意图的问题,提出行为等价基准MintEval,通过构建模块生成参考策略并比较行为,发现模型存在显著静默失败率。

中文摘要 AI 辅助

大语言模型正从生成交易信号转向编写执行这些信号的代码。第二种角色的失败模式是无声的:生成的代码运行了,回测图表绘制出来了,然而交易员所描述的风险逻辑并非实际执行的逻辑。现有的代码基准测试单元测试上的功能正确性,金融基准测试预测能力;两者均未衡量实现是否表现得像所要求的策略。我们引入了MintEval,一个基准,其中参考策略由可组合构建模块库以编程方式生成,反译为口语化的交易员指令,并由被测模型重新实现。生成的程序和参考程序在相同的市场数据和摩擦条件下逐条执行,并基于其行为而非代码相似性或利润进行比较:alpha被差分消除。MintEval v0包含800个任务,基于BTCUSDT 15分钟数据,按执行度量的状态跨度复杂度tau分层,该复杂度与描述长度解耦。低成本模型的平均ActionMatch最多为0.544,且最多精确复现0.087的任务;在200个任务的分层子集上,前沿模型(Claude Opus 5.5)达到0.889,精确复现0.575,但仍会在0.275的任务上无声失败。给定构建模块菜单,模型几乎完美地识别策略,然而79.2%的规范被正确读取的实现,在超过10%的活跃柱上出现分歧。最近一个策略生成基准的LLM评判器,逐字应用,接受了所有这些无声失败。

英文摘要

Large language models are moving from producing trading signals to writing the code that executes them. The failure mode of the second role is silent: generated code runs, a backtest plots, yet the risk logic that the trader described is not the logic being executed. Existing code benchmarks test functional correctness on unit tests and finance benchmarks test forecasting; neither measures whether an implementation behaves like the strategy that was asked for. We introduce MintEval, a benchmark in which reference strategies are generated programmatically from a library of composable building blocks, back-translated into colloquial trader instructions, and re-implemented by the model under test. Generated and reference programs are executed bar by bar on identical market data and frictions, and compared on their actions rather than on code similarity or profit: alpha is differenced away. MintEval v0 contains 800 tasks on BTCUSDT 15-minute data, stratified by an execution-measured state-span complexity tau that is decoupled from description length. Low-cost models reach a mean ActionMatch of at most 0.544 and reproduce at most 0.087 of tasks exactly; on a stratified subset of 200 tasks a frontier model (Claude Opus 5.5) reaches 0.889 and reproduces 0.575 exactly, yet still fails silently on 0.275 of tasks. Given a menu of building blocks, models identify the strategy almost perfectly, yet 79.2% of the implementations whose specification was read correctly diverge on more than 10% of active bars. The LLM judge of a recent strategy-generation benchmark, applied verbatim, accepts every one of these silent failures.

发表机构

  • Fudan University(复旦大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑