arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.39420cs.CLcs.LGq-fin.TR

QuantCode模型:面向可执行算法交易代码的语言模型特化

QuantCode Model: Specializing Language Models for Executable Algorithmic Trading Code

Alexey Chernysh, Orkhan Ekhtibarov, Dmitry Zmitrovich

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过持续预训练和监督微调特化语言模型生成可执行算法交易代码,在QuantCode-Bench基准上显著提升成功率,并识别出能力保留失败及其恢复方法。

中文摘要 AI 辅助

大型语言模型是强大的通用代码生成器,但可执行的算法交易仍是一个高要求的特化目标:模型必须将自然语言策略规范转换为特定交易框架的正确程序逻辑,在历史数据上执行,产生交易,并保持对请求的语义忠实。我们研究了两种互补的机制来特化语言模型以适应这一场景:在算法交易框架代码上进行持续预训练,以及在代理验证的请求-代码对上进行监督微调(SFT)。评估以QuantCode-Bench为中心,这是我们为Backtrader策略生成构建的400任务基准,以及一个仓库级别的类似SWE-bench的轨道。持续预训练将Qwen3.5-397B-A17B的单轮Judge Pass从41.5%提升到47.5%,将Qwen3.6-35B-A3B的从27.8%提升到33.0%。在持续预训练后应用SFT对Qwen3.6-35B-A3B产生了更大的增益,达到58.2%的Judge Pass和83.5%的成功回测;在代理评估中,它将首轮成功率从22.3%提升到58.3%,并在最多10轮后的最终成功率从47.5%提升到79.5%。仅持续预训练提高了首轮代理成功率,但将修复后的最终成功率从47.5%降至32.5%,与指令遵循能力下降一致,而SFT则同时改善了这两者。我们还发现了一个能力保留失败:领域特化降低了符合解析器的结构化工具调用能力,而针对性的恢复SFT恢复了工具调用格式,但未能恢复基础检查点的仓库级代理性能。结果表明,面向框架的预训练、验证过的SFT和显式的能力保留评估解决了领域特定可执行代码生成中的不同失败模式。

英文摘要

Large language models are strong general-purpose code generators, but executable algorithmic trading remains a demanding specialization target: a model must translate a natural-language strategy specification into correct program logic for a specialized trading framework, execute on historical data, produce trades, and remain semantically faithful to the request. We study two complementary mechanisms for specializing language models for this setting: continued pretraining on algorithmic-trading framework code and supervised fine-tuning (SFT) on agent-validated request-to-code pairs. Evaluation is centered on QuantCode-Bench, our 400-task benchmark for Backtrader strategy generation, together with a repository-level SWE-bench-like track. Continued pretraining improves single-turn Judge Pass from 41.5% to 47.5% for Qwen3.5-397B-A17B and from 27.8% to 33.0% for Qwen3.6-35B-A3B. SFT applied after continued pretraining yields a larger gain for Qwen3.6-35B-A3B, reaching 58.2% Judge Pass and 83.5% successful backtests; in agentic evaluation it raises first-turn success from 22.3% to 58.3% and final success after up to 10 turns from 47.5% to 79.5%. Continued pretraining alone improves first-turn agentic success but lowers final success after repair from 47.5% to 32.5%, consistent with degraded instruction following, whereas SFT improves both. We also identify a capability-retention failure: domain specialization degrades parser-conformant structured tool calling, and targeted recovery SFT restores tool-call formatting but not the base checkpoint's repository-level agent performance. The results show that framework-oriented pretraining, validated SFT, and explicit capability-retention evaluation address distinct failure modes in domain-specific executable code generation.

补充信息

↑