用于金融建议生成的GRPO:在CATE评估下优于商用大语言模型
GRPO for Financial Advice Generation: Outperforming Commercial LLMs under CATE Evaluation
浏览论文内容
中文总结 AI 辅助
本研究将金融建议生成建模为强化学习问题,用GRPO微调开放权重大语言模型,经独立于评判者的CATE审计,其毛利润提升约为最强商用基线的两倍,表现优于商用大语言模型。
中文摘要 AI 辅助
从业务记录生成可操作的金融建议,要求模型整合数值推理、领域知识与合理判断,同时避免可能损害业务的建议。直接监督学习存在困难:历史决策未必最优,高质量自由形式标签获取成本高昂。我们将金融建议生成建模为强化学习问题,使用分组相对策略优化(GRPO)微调一个开放权重语言模型。我们的奖励信号采用大语言模型作为评判者的评分准则,该准则从多个二元维度对每条建议的质量进行评分,并增加了防止损害的安全闸门。由于仅基于大语言模型的评估无法确认改进是否反映了真正的业务价值,而非对评判者的适配,我们补充了基于标准 doubly-robust 条件平均处理效应(CATE)估计器的独立于评判者的审计。在该观测性离线策略审计下,我们训练的大语言模型实现了约为评估中最强商用基线两倍的估计毛利润提升(0.0228 对 0.0104),同时在所有评估策略中具有最低的下行率和最小的尾部风险。值得注意的是,两种评估对基线的排名并不完全相同:未训练的基础模型在评判者准则上排名最后,但在因果审计中排名第二,这表明该审计捕捉到了评判者未捕捉到的信号。我们的结果表明,基于金融领域奖励信号的 GRPO 能够生成比商用大语言模型更实用的业务建议,且独立于评判者的因果审计是金融自然语言处理中对大语言模型作为评判者评估的有价值补充,而非对其的确认。
英文摘要
Generating actionable financial advice from business records demands that models integrate numerical reasoning, domain knowledge, and sound judgment, while avoiding recommendations that could harm the business. Direct supervision is difficult: historical decisions are not necessarily optimal, and high-quality free-form labels are expensive to obtain. We formulate financial advice generation as a reinforcement learning problem and fine-tune an open-weight language model using Group Relative Policy Optimization (GRPO). Our reward is an LLM-as-a-judge rubric that scores each recommendation across multiple binary dimensions of advice quality, augmented with a safety gate for harm prevention. Since LLM-based evaluation alone cannot confirm whether improvements reflect genuine business value rather than adaptation to the judge, we complement it with a judge-independent audit based on a standard doubly-robust Conditional Average Treatment Effect (CATE) estimator. Under this observational off-policy audit, our trained LLM achieves approximately twice the estimated gross-profit lift of the strongest evaluated commercial baseline ($0.0228$ vs.\ $0.0104$), together with the lowest downside rate and the least negative tail risk of any policy evaluated. Notably, the two evaluations do not rank the baselines identically: the untrained base model places last on the judge rubric but second on the causal audit, indicating that the audit captures a signal the judge does not. Our results demonstrate that GRPO with a finance-grounded reward signal can produce substantially more useful business recommendations than commercial LLMs, and that a judge-independent causal audit is a valuable complement to, rather than a confirmation of, LLM-as-a-judge assessment in financial NLP.
发表机构
- Intuit(英图易(金融科技公司))
机构由 AI 辅助整理,请以论文原文为准。