发表机构
Ahsanullah University of Science and Technology; Jashore University of Science and Technology; American International University - Bangladesh(阿萨努拉科技大学; 杰索尔科技大学; 孟加拉国美国国际大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出CLAIR-Fin九智能体框架,针对跨模态金融问答的声明级验证与自适应辩论,在BB-FinQA-X数据集上提升了模型忠实度,且弃权比例合理,优于相关基线方法。
AI 中文摘要
现有针对检索增强及多智能体流水线中幻觉问题的防御措施仍不完善:尽管存在模态分歧仍会信任证据,辩论仅验证汇总报告而非单个声明,且此类验证仅在起草后进行,导致智能体间错误直至最终文本才被发现。为填补这一空白,我们提出CLAIR-Fin,一个九智能体框架,将每个问题分解为在类型化金融声明账本中维护的原子声明。每个声明通过以下方式解决:非对称证据权威,其根据声明类型确定证据信任度,而非将所有模态视为同等可靠;保管链验证,在起草与对抗性审查的交接处检查依据,而非仅在流水线出口;自适应反驳循环,将有争议的声明通过对抗性辩论处理,其深度随辩论发现的内容调整;以及最终蕴含审计,搭配持续的幻觉风险指数,区分通过审查的声明与从未被争议的声明。我们在BB-FinQA-X上评估CLAIR-Fin,这是一个由孟加拉银行年度报告材料构建的500题跨模态金融评估集,按查询类型、格式和难度分层。与单次检索增强生成基线相比,它将忠实度从0.780提升至0.889,且在证据不足时对5.4%的问题弃权(不执行),而非强迫给出无依据的回答,同时在忠实度上超过HyDE和Graph-RAG等更强的检索策略基线(其忠实度≤0.874)。
英文摘要
Existing defenses against hallucination in retrieval-augmented and multi-agent pipelines remain partial: evidence is trusted despite modality disagreement, debate verifies an aggregate report rather than individual claims, and such verification occurs only after drafting, leaving inter-agent errors undetected until the final text. To close this gap, we present CLAIR-Fin, a nine-agent framework that decomposes each question into atomic claims maintained in a typed Financial Claim Ledger. Each claim is resolved through Asymmetric Evidence Authority, which conditions evidence trust on claim type rather than treating all modalities as equally reliable; Chain-of-Custody Verification, which checks grounding at the hand-off between drafting and adversarial review rather than only at the pipeline's exit; an Adaptive Rebuttal Cycle, which routes contested claims through adversarial debate whose depth scales with what that debate finds; and a terminal entailment audit paired with a continuous Hallucination Risk Index that distinguishes claims that passed scrutiny from claims never contested. We evaluate CLAIR-Fin on BB-FinQA-X, a 500-question cross-modal financial evaluation set built from Bangladesh Bank Annual Report material, stratified by query type, format, and difficulty. Relative to a single-pass retrieval-augmented generation baseline, it raises faithfulness ($0.780 \rightarrow 0.889$) while abstaining on 5.4% of questions when evidence is insufficient rather than forcing an unsupported response, and it exceeds stronger retrieval-strategy baselines such as HyDE and Graph-RAG on faithfulness ($\leq 0.874$).
CommentsAccepted at FinNLP 2026 @ EMNLP 2026