发表机构
The Hong Kong University of Science and Technology(香港科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
FinRankGRPO通过两阶段训练(思维链监督微调与组相对策略优化)将LLM用于列表式金融资产排序,以斯皮尔曼秩相关为奖励,超越传统量化与商业模型,夏普比率达0.636。
AI 中文摘要
尽管大型语言模型(LLMs)在理解非结构化金融背景方面表现出色,但它们在投资组合优化中的直接使用受到下一个词元预测与资产配置所需的列表式排序目标之间不匹配的限制。它们还难以进行精确的数值预测,导致不稳定性和算术幻觉。为了弥合这一差距,我们提出了FinRankGRPO,一个将基于LLM的投资组合构建从直接数值预测转变为金融资产列表式排序的框架。我们引入了一个两阶段的训练过程,首先在思维链推理数据上进行监督微调,然后通过我们的金融资产排序组相对策略优化,使用斯皮尔曼秩相关系数奖励,将生成的资产排序与真实市场排序对齐。第二阶段使用新颖的斯皮尔曼秩相关系数奖励,明确地将模型的生成偏好与真实市场排序对齐。实验结果表明,FinRankGRPO优于传统量化模型和最新商业模型,实现了0.636的夏普比率和0.023的斯皮尔曼相关系数。
英文摘要
While Large Language Models (LLMs) excel at understanding unstructured financial contexts, their direct use in portfolio optimization is limited by a mismatch between next-token prediction and the listwise ranking objectives required for asset allocation. They also struggle with precise numerical forecasting, leading to instability and arithmetic hallucinations. To bridge this gap, we propose FinRankGRPO, a framework that shifts LLM based portfolio construction from direct numerical prediction to listwise ranking of financial assets. We introduce a two-stage training process, supervised finetuning on Chain-of-Thought reasoning data, followed by our Financial Asset Ranking via Group Relative Policy Optimization with a Spearman rank correlation reward that aligns generated asset rankings with ground truth market orderings. The second stage uses a novel Spearman rank correlation reward to explicitly align the model's generative preferences with ground truth market orderings. Experimental results show that FinRankGRPO outperforms traditional quantitative and state-of-the-art commercial models, achieving a Sharpe ratio of 0.636 and a Spearman correlation of 0.023.