arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FinProBench:基于专业交付物衍生的角色基准评估金融AI智能体

FinProBench: Evaluating Financial AI Agents with Role-Grounded Rubrics Derived from Professional Deliverables

Ben Wang, Kang Zhou, Lifan Guo, Feng Chen, Chi Zhang

arXiv 2608.04077首次发表:更新:

AI 中文总结

该研究推出FinProBench金融AI智能体基准及RGRC角色基准构建流程,实验显示RGRC在角色专业化角色评估上优于仅提示方法,角色级复用基准可大幅减少工作量。

AI 中文摘要

评估金融AI智能体需要与实际专业工作相符的标准。现有基准方法通常从任务提示或模型输出中衍生标准,忽视了仅在从业者交付物中可见的隐性标准。我们推出FinProBench,一个面向专业金融任务的基准,以及角色基准构建(Role-Grounded Rubric Construction,RGRC),一个可复用的流程,用于从相同角色从业者产出的交付物中衍生基准。RGRC包含四个阶段:交付物收集、能力提取、基准综合与验证。其基准能捕捉隐性标准、区分质量等级,并可在同一角色的不同任务间迁移。分析前,我们按交付物类型将57种职业分为30种先验丰富的传统角色和27种先验稀疏的角色专业化角色。在所有角色中,仅提示(Prompt-only)方法在传统角色上与RGRC表现接近(89.2%对90.7%),但RGRC在角色专业化角色上显著优于它(99.1%对78.0%)。这一差异表明,当模型先验中充分包含传统标准时,提示工程可近似替代基准;而对于超出这些先验的标准,专业基准是必不可少的。FinProBench由1723份精心整理的交付物构建,涵盖57种职业、8个金融子行业和161种交付物类型,并发布了包含7个子行业20个角色的20个完整任务的初始评估集。结合异构大语言模型(LLM)评判器与角色级基准,人类交付物平均排名第一(73.7分,满分100;其他系统分别为70.3、70.2和69.6分),且所有四个系统的95%置信区间重叠,优势互补。与从头编写每个基准相比,在角色级复用基准可将每个任务的基准构建工作量估计减少6.7倍。

英文摘要

Evaluating financial AI agents requires criteria aligned with real professional work. Existing rubric methods typically derive criteria from task prompts or model outputs, overlooking tacit standards visible only in practitioner deliverables. We introduce FinProBench, a benchmark for professional financial tasks, and Role-Grounded Rubric Construction (RGRC), a reusable pipeline that derives rubrics from deliverables produced by practitioners in the same role. RGRC comprises four stages: Deliverable Collection, Competency Extraction, Rubric Synthesis, and Validation. Its rubrics capture tacit standards, distinguish quality levels, and transfer across tasks within a role. Before analysis, we classified 57 occupations by deliverable genre into 30 prior-rich conventional roles and 27 prior-sparse role-specialized roles. Across all roles, Prompt-only nearly matches RGRC for conventional roles (89.2% vs. 90.7%), but RGRC substantially outperforms it for role-specialized roles (99.1% vs. 78.0%). This split indicates that prompt engineering can approximate rubrics when conventions are well represented in model priors, while professional grounding is essential for standards beyond those priors. FinProBench is built from 1,723 curated deliverables spanning 57 occupations, 8 financial sub-industries, and 161 deliverable types, and releases an initial evaluation set of 20 complete tasks covering 20 roles in 7 sub-industries. With heterogeneous LLM judges and role-level rubrics, human deliverables rank first on average (73.7 vs. 70.3, 70.2, and 69.6 out of 100), while all four systems show overlapping 95% confidence intervals and complementary strengths. Reusing rubrics at the role level reduces estimated per-task construction effort by 6.7 times relative to authoring each rubric from scratch.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑