arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

金融服务中的自动化监管合规问答:基于领域自适应检索增强生成

Automated Regulatory Compliance Question Answering in Financial Services with Domain-Adapted Retrieval-Augmented Generation

Tobias Deußer, Abhishek Pillai, Aurelio F. Bariviera, Dhananjay Bhardwaj, Lorenz Sparrenberg, David Berghaus, Christian Bauckhage, Rafet Sifa

arXiv 2609.30009首次发表:更新:

发表机构

University of Bonn; Lamarr-Institute for Machine Learning and Artificial Intelligence; Universitat Rovira i Virgili; Fraunhofer IAIS(波恩大学; 拉马尔机器学习和人工智能研究所; 罗维拉-维尔吉利大学; 弗劳恩霍夫智能分析和信息系统研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对金融合规问答中紧凑模型幻觉问题,提出基于LegalBERT的三阶段检索器与RAFT-LoRA生成器的领域自适应RAG流水线,在ObliQA上显著提升检索与答案质量,但发现生成增益未证明接地性且跨域迁移受限。

AI 中文摘要

金融机构在密集且频繁修订的规则手册下运营,正确回答合规问题不仅需要流畅性,还需要在权威文本中有可验证的依据。大型语言模型对此任务具有吸引力,但企业实际可在本地部署的模型是紧凑型模型,而紧凑型模型会产生义务幻觉。我们研究精心设计的领域自适应检索增强生成流水线是否能弥合这一差距。我们的检索器在LegalBERT之上分三个阶段构建:蕴含调优,将问题-段落匹配重构为前提-假设重建;使用批内负样本进行对比调优;以及与BM25的分数级融合。我们的生成器是一个紧凑模型(2B-12B参数),以4位量化方式服务,可通过提示或通过LoRA进行检索感知微调(RAFT)来适配。在ObliQA(一个基于阿布扎比全球市场规则手册构建的问答基准)上,分阶段检索器将Recall@10从0.256提升至0.774,并优于BM25(0.678)和我们测试过的最强通用密集编码器E5-large-v2(0.758)。随后,RAFT-LoRA为我们能适配的每个模型提高了综合RePASs答案质量分数,其中对最弱模型的提升最大。然而,适配后的模型无法迁移到澳大利亚判例法问题,而一个完全不接收段落的闭卷模型在RePASs上得分与完整流水线相差0.011以内,同时生成的答案不引用任何内容并错误陈述义务。因此,检索增益是直接测量的,生成增益是RePASs上的增益而非已证明的接地性,而接地性本身需要RePASs未提供的评估协议。

英文摘要

Financial institutions operate under dense, frequently amended rulebooks, and answering a compliance question correctly requires not only fluency but verifiable grounding in the authoritative text. Large language models are attractive for this task, yet the models that firms can realistically deploy on-premise are compact ones, and compact models hallucinate obligations. We study whether a carefully domain-adapted retrieval-augmented generation pipeline closes that gap. Our retriever is built in three stages on top of LegalBERT: entailment tuning that recasts question--passage matching as premise--hypothesis reconstruction, contrastive tuning with in-batch negatives, and score-level fusion with BM25. Our generator is a compact model (2B--12B parameters) served under 4-bit quantization, either prompted or adapted with retrieval-aware fine-tuning (RAFT) through LoRA. On ObliQA, a question-answering benchmark built from the Abu Dhabi Global Market rulebooks, the staged retriever raises Recall@10 from 0.256 to 0.774 and outperforms BM25 (0.678) and E5-large-v2 (0.758), the strongest general-purpose dense encoder we tested. RAFT-LoRA then improves the composite RePASs answer-quality score for every model we could adapt, with the largest gain on the weakest one. However, the adapted models do not transfer to Australian case-law questions, and a closed-book model that receives no passages at all scores within 0.011 RePASs of the full pipeline while producing answers that cite nothing and misstate obligations. The retrieval gain is therefore measured directly, the generation gain is a gain in RePASs rather than demonstrated grounding, and grounding itself requires an evaluation protocol that RePASs does not provide.

CommentsCurrently under review

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑