FinReportBench:衡量与提升机构级财务报告生成能力
FinReportBench: Measuring and Improving Institution-Grade Financial Report Generation
浏览论文内容
中文总结 AI 辅助
本文提出FinReportBench基准,发现大型语言模型生成机构级财务报告的瓶颈在于报告身份认同与机构完整性,通过基准引导的技能蒸馏可显著提升模型相关指标。
中文摘要 AI 辅助
大型语言模型能够生成流畅的财务分析,但仅流畅性不足以判断报告是否适合机构交付。本文介绍FinReportBench,这是一个基于专家意见的基准,用于衡量和提升机构级财务报告生成能力。专家评审发现报告存在身份认同、机构组成、来源学科及可视化交付方面的反复出现的缺陷。我们通过专家偏序关系、多模态证据及决策边界审计,得出包含35个条目的评分标准,覆盖可交付性、报告身份认同及机构完整性。我们从10000条平衡的中英双语财务研究源记录出发,整理出244个双语任务,涉及三类研究对象和两类输入层级,每个任务区分公开查询、重构的研究轨迹及隐藏的源数据包。三类独立评审组以接近天花板的比率复现了专家偏序关系,表明有界的可观测标准支持可靠评估。在九类模型中,基础可交付性已接近饱和,而报告身份认同和机构完整性仍是主要瓶颈;模型间最大差距在于生成轨迹控制、信息密度及数据规范,而非基础报告框架。随后,我们使用基准引导的技能蒸馏,将反复出现的失败转化为可复用的生成及自评审约束。在五类模型中,经优化的技能使平均G1提升33.85分,平均G2提升13.83分,且每对模型均保留G0。代码及基准 artifacts 可在this https URL获取。
英文摘要
Large language models can produce fluent financial analysis, but fluency alone does not establish whether a report is suitable for institutional delivery. We introduce FinReportBench, an expert-grounded benchmark for measuring and improving institution-grade financial report generation. Expert review reveals recurring gaps in report identity, institutional components, source discipline, and visual delivery. We derive a 35-item rubric through expert partial orders, multimodal evidence, and audits of decision boundaries, covering deliverability, report identity, and institutional completeness. Starting from 10,000 balanced Chinese and English financial-research source records, we curate 244 bilingual tasks across three research objects and two input tiers. Each task separates the public query, reconstructed research trajectory, and hidden source packet. Three independent judge families reproduce the expert partial order at near-ceiling rates, showing that bounded, observable criteria support reliable evaluation. Across nine model families, basic deliverability is nearly saturated, while report identity and institutional completeness remain the primary bottlenecks. The largest cross-model gaps concern generation-trace control, information density, and data discipline rather than basic report framing. We then use benchmark-guided skill distillation to turn recurrent failures into reusable generation and self-review constraints. Across five model families, the evolved skill improves mean G1 by 33.85 points and mean G2 by 13.83 points over paired no-skill runs while preserving G0 for every pair. Code and benchmark artifacts are available at https://github.com/MisterBrookT/finreportbench.