发表机构
Ant International; Xiamen University; Shanghai University; Tongji University(蚂蚁集团; 厦门大学; 上海大学; 同济大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究推出面向金融风险审查的中文基准FinRiskAtlas,从操作执行与证据状态控制两维度评估金融大语言模型,发现通用金融能力评分无法完全反映模型在专业工作流程中的可靠性,需采用决策对齐的评估单元。
AI 中文摘要
将大语言模型部署到专业金融审查场景,仅衡量通用金融能力是不够的:模型必须执行工作流程所需的特定审查操作,并确定现有证据是否足以做出可辩护的决策。现有的金融基准涵盖知识、推理、合规性和专业任务,但它们的评估单元通常围绕数据集或任务形式组织,而非部署系统所支持的决策。我们推出FinRiskAtlas,这是一个中文基准,从两个互补维度评估金融大语言模型:固定证据状态下的操作执行,以及审查条件演变下的证据状态控制。静态基准包含53个任务系列共9742个实例,其中包括42个领域知识系列和11个由明确评估合约定义的下游审查操作。FinRisk-Ask通过离线重放104条去标识化专业轨迹中的680个行动前状态来扩展该框架,在推理过程中 withheld未来证据,仅用其构建专家验证的证据目标。在33种模型配置中,操作级评估产生非冗余排名(下游操作间的平均成对斯皮尔曼相关系数为0.42),而基于知识的筛选在单个操作上可能产生高达18.01分的遗憾值。FinRisk-Ak进一步显示,更频繁进入Ask分支并不一定能改善请求定位或端到端证据获取。这些结果表明,广泛的金融能力评分无法完全捕捉模型在专业工作流程中可靠的场景,推动了与部署系统必须支持的决策和证据状态对齐的评估单元的发展。
英文摘要
Deploying large language models for professional financial review requires more than measuring general financial competence: models must perform the specific review operation required by a workflow and determine whether available evidence is sufficient for a defensible decision. Existing financial benchmarks cover knowledge, reasoning, compliance, and professional tasks, but their evaluation units are often organized around datasets or task formulations rather than the decisions that deployed systems support. We introduce FinRiskAtlas, a Chinese-language benchmark that evaluates financial LLMs along two complementary dimensions: operation execution under fixed evidence states and evidence-state control under evolving review conditions. The static benchmark contains 9,742 instances across 53 task families, including 42 Domain Knowledge families and eleven downstream review operations defined by explicit evaluation contracts. FinRisk-Ask extends this framework through offline replay of 680 pre-action states from 104 de-identified professional trajectories, withholding future evidence during inference and using it only to construct expert-verified evidence targets. Across 33 model configurations, operation-level evaluation yields non-redundant rankings (mean pairwise Spearman correlation 0.42 across downstream operations), and knowledge-based shortlisting can incur up to 18.01 points of regret on individual operations. FinRisk-Ask further shows that entering the Ask branch more frequently does not necessarily improve request targeting or end-to-end evidence acquisition. These results show that broad financial capability scores do not fully capture where models are reliable in professional workflows, motivating evaluation units aligned with the decisions and evidence states that deployed systems must support.