发表机构
Princeton University; InclusionAI; Harvard University; Stanford University(普林斯顿大学; InclusionAI; 哈佛大学; 斯坦福大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过后训练(SFT+PPO)提升语言模型预测股票价格的能力,提出AURA-4B模型,在方向-幅度分数上从20.94翻倍至43.31,达到前沿水平,证明后训练可显著改善金融预测。
AI 中文摘要
后训练已被证明能显著提升语言模型在具有可验证结果的任务上的表现,包括数学推理、软件工程和计算机使用。然而,同样的方法能否改善金融市场预测则远不那么明确。与具有可验证结果的任务相比,不仅已实现收益存在噪声,而且即使是构成有效预测所需的相关信息集,在事先也并非显而易见:模型必须决定收集哪些观测数据,然后在结果揭晓之前承诺一个数值判断。我们在一个按时间顺序排列的股票价格沙盒环境中研究这个问题,其中语言模型收集价格、成交量、相对表现和市场背景证据,并预测未来收益。我们使用监督微调(SFT)对Qwen3-4B进行后训练,基于工具使用演示,然后使用近端策略优化(PPO),以预测分数与已实现收益对比给出的终端奖励。由此得到的AURA-4B将起始方向-幅度分数从20.94提升至43.31,翻了一倍多,并且在该基准上可与前沿语言模型相媲美。条件幅度一致性从33.3上升至66.2,而方向准确率从62.9变化至65.4。SFT扩展了工具使用,PPO进一步增加了排名和市场背景查询的份额。这些结果表明,在这个以结果为导向的基准上,后训练可以显著提升金融预测性能,同时改变模型调查市场的方式。
英文摘要
Post-training has been shown to significantly improve language models' performance on tasks with verifiable outcomes, including mathematical reasoning, software engineering, and computer use. However, whether the same approach can improve forecasting in financial markets is much less clear. Compared with tasks with verifiable outcomes, not only are realized returns noisy, but even what constitutes a relevant information set for making effective predictions is not obvious a priori: the model must decide which observations to gather and then commit to a numerical judgment before the outcome is known. We study this question in a chronological stock-price sandbox, where a language model gathers price, volume, relative-performance, and market-context evidence and predicts a future return. We post-train Qwen3-4B with supervised fine-tuning (SFT) on tool-use demonstrations, then proximal policy optimization (PPO) with a terminal reward given by the forecast score against the realized return. The resulting AURA-4B more than doubles the starting direction--magnitude score, from 20.94 to 43.31, and is comparable to frontier language models on this benchmark. Conditional magnitude agreement rises from 33.3 to 66.2, while directional accuracy changes from 62.9 to 65.4. SFT expands tool use, and PPO further increases the share of ranking and market-context queries. These results show that post-training can substantially improve financial forecasting performance, together with changes in how the model investigates the market, on this outcome-selected benchmark.
Comments18 pages, 4 figures