arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FinancialAuditBench:基于真实世界先验的差分隐私基准构建

FinancialAuditBench: Benchmark Construction under Differential Privacy Using Real-World Priors

Jerry Huang, Sarvesh Babu, Matt Van Buren, Alexander Wang, Pranav Pillai, Arush Jain, James P. Burton, Julia Hockenmaier

arXiv 2609.32835首次发表:更新:

AI 中文总结

针对AI智能体在金融审计中部署的可靠性测量问题,提出FinancialAuditBench基准,利用差分隐私统计和审计专业知识生成合成任务,评估显示智能体可完成多数员工级任务但仍有错误。

AI 中文摘要

随着AI智能体在金融服务行业中被广泛采用,仔细的测量对于理解它们可以在哪些地方可靠部署以及哪些地方仍然需要监督和专业审查至关重要。然而,这种测量受到对专有或隐私敏感数据访问受限的制约。因此,现有基准通常依赖于公开可用的数据、由人类和/或LLM撰写的任务或简化设置。我们引入了FinancialAuditBench,一个用于在财务报表审计任务上评估智能体的基准,以及一个系统生成合成业务的框架。我们的任务生成框架利用来自历史审计的差分隐私聚合统计以及通过超过1100小时的基准开发和审查贡献的审计专业知识。FinancialAuditBench包含90个任务,涵盖六个合成审计业务中的工作底稿完成和审查,每个业务平均包含179个文件。对十一个前沿模型的评估表明,虽然智能体能够很好地完成大部分员工级别的审计任务,但它们有时会执行不适当的程序或生成不正确的文档。除了财务审计之外,我们的框架还为在隐私敏感领域系统生成用于模型评估和训练的合成任务提供了一种方法。

英文摘要

As AI agents are becoming widely adopted in the financial services industry, careful measurement is essential to understand where they can be reliably deployed and where oversight and professional review remain necessary. Such measurement, however, is constrained by limited access to proprietary or privacy-sensitive data. Existing benchmarks therefore often rely on publicly available data, human- and/or LLM-authored tasks, or simplified settings. We introduce FinancialAuditBench, a benchmark for evaluating agents on financial statement audit tasks, along with a framework for systematically generating synthetic engagements. Our task generation framework leverages differentially private aggregate statistics from historical audits along with audit expertise contributed through over 1,100 hours of benchmark development and review. FinancialAuditBench consists of 90 tasks spanning workpaper completion and review across six synthetic audit engagements, each containing an average of 179 files. Evaluation on eleven frontier models shows that while agents complete substantial portions of staff-level audit tasks well, they sometimes perform inappropriate procedures or produce incorrect documentation. Beyond financial auditing, our framework offers an approach for systematically generating synthetic tasks for model evaluation and training in privacy-sensitive domains.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑