发表机构
University of California, Berkeley; California Polytechnic State University; Hofstra University(加州大学伯克利分校; 加州州立理工大学; 霍夫斯特拉大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对专家级金融问答的合规与防幻觉需求,提出自改进RAG框架,通过三个专门智能体及反馈自校正实现,在FinanceBench数据集上获86%神谕引导准确率,为金融监管应用提供可解释方案。
AI 中文摘要
专家级金融问答既需要基于事实的验证以捕捉数值幻觉,又需要审计追踪以满足监管合规要求,这是标准的单次检索增强生成(RAG)系统所欠缺的特性。我们通过自改进RAG(Self-Improving RAG)框架向该目标迈进,该框架将文档问答分解为三个专门智能体(检索、推理和裁决),由编排器通过反馈驱动的自校正进行协调。当裁决智能体对答案的评分低于动态阈值时,系统会触发重试并采用升级策略:更广泛的检索、更谨慎的提示以及更宽松的接受标准。我们在FinanceBench(美国证券交易委员会文件问答)数据集上进行评估,自改进RAG实现了86%的神谕引导准确率(衡量与标准答案的一致性),并达到36.4%的拉撒路率(Lazarus Rate),通过针对性重试恢复了近十分之四的初始错误答案。一项关键发现是,带有裁决驱动重试的固定检索管道无需动态路由即可实现优异结果,并具备完全可解释性;每个决策都记录了置信度分数,为受监管的金融应用提供了所需的审计追踪。
英文摘要
Expert-level financial question answering requires both grounded verification to catch numeric hallucinations and audit trails for regulatory compliance, attributes that standard single-pass RAG systems lack. We take a step toward this goal with Self-Improving RAG, a framework that decomposes document QA into three specialized agents (Retrieval, Reasoning, and Judge) coordinated by an orchestrator with feedback-driven self-correction. When the Judge Agent scores an answer below a dynamic threshold, the system triggers retry with escalated strategies: broader retrieval, more careful prompting, and relaxed acceptance criteria. We evaluate on FinanceBench (SEC filing QA), where Self-Improving RAG achieves 86% oracle-guided accuracy (measuring agreement with gold answers) with a 36.4% Lazarus Rate, recovering nearly 4 in 10 initially incorrect answers through targeted retry. A key finding is that a fixed retrieval pipeline with judge-driven retry achieves strong results without dynamic routing, providing full interpretability. Every decision is logged with confidence scores, enabling the audit trails required for regulated financial applications.
Comments17 pages, 2 figures. Accepted at the ICLR 2026 Workshop on Advances in Financial AI