arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.07603cs.AI

FinCUABuild:智能体能否为动态金融计算机使用构建可靠基准?

FinCUABuild: Can Agents Build Reliable Benchmarks for Dynamic Financial Computer Use?

  • School of Artificial Intelligence, Wuhan University(武汉大学人工智能学院)
  • MBZUAI(穆罕默德·本·扎耶德人工智能大学)
  • Northeastern University(东北大学)
  • Zhongguancun Academy(中关村学院)

机构由 AI 辅助整理,请以论文原文为准。

Jingpu Yang, Fengxian Ji, Jinri Guo, Tianhao Li, Qian Jiang, Fan Zhang, Min Peng, Qianqian Xie, Preslav Nakov, Zhuohan Xie

AI总结:

本文提出FinCUABuildBench基准和FinCUABuildAgent多智能体系统,用于自动构建金融CUA评估任务,显著提升严格资格率至31.3%,验证了智能体自主构建可靠基准的可行性。

AI中文摘要:

金融场景多样且复杂,涵盖不同的数据条件、工具配置和工作流程。然而,现有的CUA(计算机使用智能体)评估任务大多依赖人工构建,限制了真实金融场景的可扩展覆盖。那么,智能体能否自主构建多样化的金融CUA评估任务?评估这一能力面临三个关键挑战:构建请求的场景覆盖、不同构建方法之间的公平比较,以及生成任务质量的可靠评估。为解决这些问题,我们引入了FinCUABuildBench,一个用于评估金融CUA任务构建的基准,其特点包括:(i)576个构建请求,覆盖24个金融工作流程和三种类型的运行时变化;(ii)标准化的输入、预算和输出规范;(iii)基于执行测试和质量检查的任务资格验证机制。我们进一步引入了FinCUABuildAgent,一个用于自动构建动态金融CUA评估任务的多智能体系统。它由三个模块组成,共同构建任务、环境和验证器。在FinCUABuildBench上,在相同的模型骨干下,现有的基于智能体的构建方法的严格资格率仅为1.3%至8.3%,而FinCUABuildAgent达到了31.3%。下游评估进一步表明,所构建的任务能够有效区分CUA任务执行能力。这些结果表明,智能体能够自主构建具有有意义评估价值的金融CUA任务,为金融场景中更广泛的评估覆盖提供了一条实用路径。代码:此https URL

英文摘要:

Financial scenarios are diverse and complex, spanning varying data conditions, tool configurations, and workflows. Yet existing CUA, Computer-Using Agent, evaluation tasks remain largely manually constructed, limiting scalable coverage of real-world financial scenarios. Then, can agents autonomously construct diverse CUA evaluation tasks for financial scenarios? Evaluating this capability poses three key challenges: scenario coverage of construction requests, fair comparison across construction methods, and reliable assessment of generated task quality. To solve these, we introduce FinCUABuildBench, a benchmark for evaluating financial CUA task construction, featuring: (i) 576 construction requests covering 24 financial workflows and three types of runtime variation; (ii) standardized input, budget, and output specifications; and (iii) a task qualification mechanism based on execution tests and quality checks. We further introduce FinCUABuildAgent, a multi-agent system for automatically constructing dynamic financial CUA evaluation tasks. It consists of three modules that jointly construct tasks, environments, and validators. On FinCUABuildBench, under the same model backbone, existing agent-based construction methods achieve strict qualification rates of only 1.3-8.3%, while FinCUABuildAgent reaches 31.3%. Downstream evaluations further show that the constructed tasks can effectively differentiate CUA task-execution capabilities. These results demonstrate that agents can autonomously construct financial CUA tasks with meaningful evaluation value, offering a practical path toward broader evaluation coverage in financial scenarios. Code: https://github.com/FengxianJi/FinCUABuild

补充信息

↑