arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.35744cs.AIq-fin.CP

FinAutoRubric:专家引导的金融研究智能体自动评分标准生成

FinAutoRubric: Expert-Guided Automatic Rubric Generation for Evaluating Financial Research Agents

Hoyoung Lee, Suyeol Yun, Jack Haverty, Yunju Cho, Meesong Kim, Daekyung Park, Sumin Kim, Jihoon Kwon, Jasmine Jia Geng, Andrew Chin, Yin Luo, Edward Tong, Yu Yu… 展开作者

Hoyoung Lee, Suyeol Yun, Jack Haverty, Yunju Cho, Meesong Kim, Daekyung Park, Sumin Kim, Jihoon Kwon, Jasmine Jia Geng, Andrew Chin, Yin Luo, Edward Tong, Yu Yu, Zach Golkhou, Minkyu Kim, Igor Halperin, Young Cha, Alejandro Lopez-Lira, Chanyeol Choi, Yongjae Lee

首次发表
浏览论文内容

中文总结 AI 辅助

FinAutoRubric通过专家指导的自动评分标准生成,提升金融研究智能体评估的准确性与可扩展性,其评分与专家及人工高度一致。

中文摘要 AI 辅助

评估金融研究智能体需要反映专家标准并固定信息截止日期时正确值的评分标准。专家评审的金融基准依赖于固定的、逐项的评分标准,这些标准扩展成本高昂,且无法编码各机构自身的标准。在FinAutoRubric中,专家指定可复用的评估指导,而智能体和代码则执行查询特定的评分标准生成、审查和验证。该专家指导以提示词和代码执行的规则形式管理每个智能体,可复用标准的任务库将其跨任务传递。在遵循专家指导的长时间循环中,编写智能体研究每个期望值,审查智能体验证该值,失败则升级至人工处理。在三个专家撰写的金融基准上,其评分标准在追踪专家评分方面与最强的评估生成器相当,同时为更多标准陈述了专家评分标准的期望值,其评分与人工评分一致,内部分析师在盲审中更偏好它们。发布的包含100个查询的FinAutoRubric基准,基于内部分析师跨78个任务和八类资产的关键问题构建,表明早期模型代生成的评分标准仍为后期模型留有提升空间。

英文摘要

Evaluating finance research agents requires rubrics that reflect expert standards and fix the values correct as of an information cutoff. Expert-reviewed finance benchmarks rely on fixed, per-item rubrics, which are costly to extend and cannot encode each institution's own standard. In FinAutoRubric, experts specify reusable evaluation guidance, while agents and code carry out query-specific rubric generation, review, and validation. This expert guidance governs every agent, as prompts and as rules that code enforces, and a Task Bank of reusable criteria carries it across tasks. In long-horizon loops that follow the expert guidance, a writer agent researches every expected value and a reviewer agent verifies it, and failures escalate to a human. On three expert-authored finance benchmarks, its rubrics track expert scoring as closely as the strongest evaluated generator while stating the expert rubric's expected value for more criteria, their scores agree with human grading, and in-house analysts prefer them in a blind review. The released 100-query FinAutoRubric Benchmark, built from in-house analysts' key questions across 78 tasks and eight asset classes, shows that rubrics from an earlier model generation still leave headroom for a later one.

发表机构

  • LinqAlpha
  • UNIST(蔚山科学技术院)
  • MassMutual Life Insurance(万通互惠人寿保险)
  • AllianceBernstein(联博)
  • Wolfe Research(沃尔夫研究公司)
  • Google(谷歌)
  • BlackRock(贝莱德)
  • J.P. Morgan Chase(摩根大通)
  • State Street Corporation(道富集团)
  • Fidelity Investments(富达投资)
  • Blackstone(黑石集团)
  • University of Florida(佛罗里达大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑