发表机构
Microsoft Corporation(微软公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对智能代理在企业金融领域应用,提出FORCE-Bench基准测试,含251个专家注释查询,从八个维度评估三种任务类型,在特定设置下评估通用和专用代理,结果显示专用代理更可靠,相关资源开源助力企业金融环境应用。
AI 中文摘要
大型语言模型的进展加速了智能代理系统在运营金融中的部署。现有基准测试侧重于衡量一般能力、指令遵循或安全性,很少直接针对智能代理系统正在部署以自动化的运营金融工作流程。金融专业人员要求代理不仅提供事实准确且有充分依据的信息,还要确保信息可验证并始终符合运营金融领域的规则和约束。我们引入FORCE-Bench,它包含251个专家注释的查询,并使用基于评分标准的框架进行评估,该框架针对运营金融领域的要求进行校准,涵盖准确性、引用、清晰度、深度、依据性、时效性、相关性和结构八个维度。FORCE-Bench在三种任务类型上评估智能代理系统:财务义务研究、金融实体绩效研究和业务简报生成。为反映实际部署条件,我们在通用工具访问和延迟受限设置下评估了我们专门构建的代理以及通用智能代理系统。结果表明,通用智能代理系统在运营约束下不能始终满足金融领域的质量要求,而专门为Microsoft 365 Copilot构建的金融代理在各维度上更可靠。我们将数据集、评分标准、工具和分析代码作为开源发布,以支持可重复的比较并适应其他企业金融环境。
英文摘要
Recent advances in large language models have accelerated deployment of agentic systems in operational finance. Existing benchmarks emphasize measuring general capabilities, instruction following, or safety, but few directly address the operational finance workflows that agentic systems are now being deployed to automate. Finance professionals require agents to not only provide factually sound and properly grounded information, but also ensure that this information is verifiable and consistently adheres to rules and constraints of the operational finance domain. We introduce FORCE-Bench, which contains 251 expert-annotated queries and evaluates responses using a rubric-based framework calibrated to the requirements of the operational finance domain, across eight dimensions: accuracy, citations, clarity, depth, groundedness, recency, relevance, and structure. FORCE-Bench assesses agentic systems on three task types: financial obligation research (querying ERP systems for accounts receivable and payable data), financial entity performance research (answering time-bound questions from public filings and market data), and business brief generation (synthesising multi-source company intelligence reports). To reflect real deployment conditions, we evaluate our purpose-built agent, as well as the general-purpose agentic systems, under common tool access and latency-bounded settings. Results show that general-purpose agentic systems do not consistently meet finance-domain quality requirements under operational constraints, while the purpose-built Finance Agent for Microsoft 365 Copilot is more reliable across dimensions. We release the dataset, rubrics, harness, and analysis code as open-source to support reproducible comparison and adaptation to other enterprise finance environments.
Comments22 pages, 10 figures