arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.19259cs.LGcs.AI

财务报表欺诈检测中的泛化基准测试:稳健评估与新任务

Benchmarking Generalization in Financial Statement Fraud Detection: robust evaluation and novel tasks

Guy Stephane Waffo Dzuyo, Gaël Guibon, Christophe Cerisara, Luis Belmar-Letelier

首次发表
浏览论文内容

中文总结 AI 辅助

针对财务报表欺诈检测中现有方法性能估计不现实的问题,提出利用大语言模型整合结构化与非结构化数据的框架,通过新基准任务评估,构建公开数据集,该方法在新任务中性能最佳,凸显文本数据和稳健评估的价值。

中文摘要 AI 辅助

财务报表欺诈检测(FSFD)对市场诚信至关重要,但面临着欺诈手段日益复杂以及财务报告中文本数据未充分利用的挑战。现有方法常依赖随机数据分割,导致性能估计过于乐观,无法反映对新公司或未来时期的实际泛化能力。为解决这一问题,我们提出了一个稳健的FSFD框架,利用大语言模型(LLMs)整合结构化财务数据和财务报告中的非结构化文本信息。通过一个名为公司隔离FSFD(CI-FSFD)的新颖且具有挑战性的基准任务,我们提供了更现实的评估。我们构建并公开了一个综合的美国公司数据集,结合了财务报表、摘要式MD&A文本和欺诈标签。我们的方法在具有挑战性的CI-FSFD任务上取得了最佳性能,证明了文本数据和稳健评估对可靠财务欺诈检测的关键价值。

英文摘要

Financial statement fraud detection (FSFD) is crucial for market integrity but faces challenges from increasingly sophisticated schemes and under-utilized textual data in financial reports. Existing methods often rely on random data splits, leading to overoptimistic performance estimates that do not reflect real-world generalization to new companies or future periods. To address this recurring problem with the state of the art, we propose a robust FSFD framework leveraging Large Language Models (LLMs) to integrate both structured financial data and unstructured textual information from financial reports. We provide a more realistic evaluation through a novel and challenging benchmark task called Company-Isolated FSFD (CI-FSFD). We construct and make publicly available a comprehensive U.S. company dataset combining financial statements, summarized MD&A text, and fraud labels. Our approach achieves the best performance on the challenging CI-FSFD task, demonstrating the critical value of textual data and robust evaluation for reliable financial fraud detection.

补充信息

↑