发表机构
The Hong Kong University of Science and Technology (Guangzhou); Nanyang Technological University(香港科技大学(广州); 南洋理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对预测市场策略难以端到端运营化的问题,提出AlphaOpsBench基准,评估从经济假设到可执行程序的生成,发现严格有效性罕见,可回放性远弱于来源忠实运营化。
AI 中文摘要
大型语言模型日益生成量化交易策略,然而现有基准测试假设标准化资产、数值特征或可直接编译的策略表示——这些假设在预测市场策略中不成立,因为一个粗略的想法可能使交易结果、因果信息来源、信号定义、阈值、仓位规模、订单策略、退出和结算行为未明确。我们引入AlphaOpsBench,它评估从基于来源的经济假设到可审计可执行程序的端到端运营化,涵盖581条保留来源的策略记录和一个生命周期规模的Polymarket数据集,包含128万个二元市场、1.836亿条清洗后的执行记录、结算证据和限价订单簿历史,比较直接生成与分阶段设计-然后-编码协议。在一项修正的独立生成研究中,涵盖36个受控任务和24个预注册的真实策略,严格的端到端有效性仍然罕见:直接和分阶段在受控队列上分别获得35/180和20/180的规范通过,在真实队列上无确认通过,且重复生成在模型拥有的经济选择上差异显著。相比之下,783,655次计划的历史回放中有775,725次完成,表明可回放性远弱于来源忠实的运营化。财务结果取决于声明的执行模型和可用的历史证据,费用和流动性实验表明执行成本改变后续交易路径,而非仅作为事后扣除。因此,AlphaOpsBench在预测市场中为基于LLM的量化研究提供了一个证据感知的基准,区分策略保真度、行为有效性、历史可执行性和财务表现。
英文摘要
Large language models increasingly generate quantitative trading strategies, yet existing benchmarks assume standardized assets, numerical features, or directly compilable strategy representations---assumptions that prediction-market strategies violate, since a coarse idea may leave the traded outcome, causal information source, signal definition, threshold, sizing, order policy, exit, and settlement behavior unspecified. We introduce \textsc{AlphaOpsBench}, which evaluates end-to-end operationalization from source-grounded economic hypotheses to auditable executable programs over 581 source-preserving strategy records and a lifecycle-scale Polymarket dataset with 1.28 million binary markets, 183.6 million cleaned executions, settlement evidence, and limit-order-book history, comparing Direct generation against a Staged design-then-code protocol. In a corrected independent-generation study over 36 controlled tasks and 24 preregistered real strategies, strict end-to-end validity remains rare: Direct and Staged obtain 35/180 and 20/180 canonical passes on the controlled cohort and no confirmed pass on the real cohort, and repeated generations vary substantially in model-owned economic choices. By contrast, 775,725 of 783,655 scheduled historical replays complete, showing that replayability is a far weaker property than source-faithful operationalization. Financial outcomes depend on the declared execution model and available historical evidence, and fee and liquidity experiments show that execution costs alter subsequent trading paths rather than acting only as ex-post deductions. \textsc{AlphaOpsBench} thus separates strategy fidelity, behavioral validity, historical executability, and financial performance in an evidence-aware benchmark for LLM-based quantitative research in prediction markets.