发表机构
College of Computer Science and Artificial Intelligence, Fudan University; School of Data Science, Fudan University; Ant Group; School of Information and School of Smart Governance, Renmin University of China(复旦大学计算机科学与人工智能学院; 复旦大学数据科学学院; 蚂蚁集团; 中国人民大学信息学院与智慧治理学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文构建首个金融错误检测基准FinED-Bench,评估发现当前LLMs难胜任高复杂度金融文档错误检测任务,监督微调可提升弱模型性能。
AI 中文摘要
确保金融文档的准确性对经济分析、合规监管及企业决策至关重要。多项研究表明,大语言模型(LLMs)在股价走势预测、金融分析等诸多金融任务中表现出色,但仍有一项关键任务未被探索:LLMs识别金融文档错误的能力。本文提出首个公开的金融错误检测基准测试FinED-Bench,涵盖三个认知复杂度级别,包含9种真实金融场景、900余份现有语言模型未见过的2025年报告文档。本文详细介绍该基准测试的构建流程,并在需要金融领域知识与推理能力的任务上,评估了GPT-4o、Qwen3-14B等多款先进LLMs。实验结果显示,当前LLMs在该任务上仍存在困难,尤其在高复杂度案例中表现更差;此外,监督微调可显著提升性能较弱的LLMs在该任务上的表现。相关数据与代码可在指定网址获取。
英文摘要
Ensuring the accuracy of financial documents is critical for economic analysis, regulatory compliance, and corporate decision-making. Several studies have shown that Large Language Models (LLMs) perform well in many financial tasks, such as stock price movements and financial analytics. However, a critical task remains unexplored: the ability of LLMs to identify errors in financial documents. In this paper, we introduce \textbf{FinED-Bench}, the first publicly \textbf{Bench}mark for \textbf{Fin}ancial \textbf{E}rror \textbf{D}etection across three levels of cognitive complexity. FinED-Bench covers nine real-world financial scenarios, and includes over 900 documents reported in 2025 that are unseen by existing language models. We detail the benchmark construction process and evaluate several advanced LLMs (e.g., GPT-4o, Qwen3-14B) on this tasks, which requires both financial domain knowledge and reasoning capabilities. Experimental results show that current LLMs still struggle with this task, especially in high-complexity cases. Besides, supervised fine-tuning can significantly improve the performance of weaker LLMs on this task. Our data and code are available at https://github.com/hedyHe/FinED-Bench.