arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FinRegQA-EU:基于腐败偏好数据的欧盟金融监管问答

FinRegQA-EU: Corruption-Based Preference Data for Grounded EU Financial Regulatory Question Answering

Aulia Kharis Rakhmasari, Fan Yu, Alexander Hoyle, Elliot Ash

arXiv 2609.36856首次发表:更新:

发表机构

ETH Zürich(苏黎世联邦理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对欧洲金融监管问答,提出基于官方语料库的偏好数据流水线,发现SFT虽得分高但引用虚构严重,DPO/GRPO更优,并揭示LLM-as-a-Judge评估盲点。

AI 中文摘要

大型语言模型(LLMs)在处理特定区域的事实性知识方面存在困难,尤其是在金融监管领域。尽管诸如CFinBench等基准测试在其他地区提供了广泛的金融知识覆盖,但欧洲金融监管领域尚无类似资源。我们通过提出一个端到端的流水线来弥合这一差距,用于评估和改进LLMs在欧洲金融监管问答上的表现。我们的数据集基于欧洲银行管理局(EBA)和欧洲证券与市场管理局(ESMA)的官方问答语料库。我们采用点式LLM-as-a-Judge协议,使用来自三个不同模型家族的评判员评估候选答案,仅保留一致同意的配对,并通过监管失败模式分类法(包括法律互换、条款互换、虚构引用和虚构文本)来强化被拒绝的答案。通过对生成的偏好对进行微调,揭示了我们研究结果的核心分歧:监督微调(SFT)在所有方法中获得了最高的评判员评分,但基于规则的引用F1分数却最低。SFT生成了更长的答案,引用数量是其他方法的三倍,但其中大多数引用缺乏支持。DPO和GRPO提供了更好的权衡,答案更简洁,且在微调模型中获得了最高的引用F1分数。标准的LLM-as-a-Judge评估将引用密度视为接地性的证据,但无法检测这些引用是否被虚构。因此,我们发现在答案必须可验证的场景中存在的盲点。我们在提供的URL发布了基准测试和评估栈。

英文摘要

Large Language Models (LLMs) struggle with region-specific factual knowledge, particularly in financial regulation. While benchmarks such as CFinBench, provide broad coverage of financial knowledge in other regions, no comparable resource exists for European financial regulation. We close this gap by proposing an end-to-end pipeline for evaluating and improving LLMs on European financial regulatory question answering. Our dataset is grounded in the official Q&A corpora of the European Banking Authority (EBA) and the European Securities and Markets Authority (ESMA). We evaluate candidate answers with a pointwise LLM-as-a-Judge protocol using three judges from distinct model families, retain only unanimously judged pairs, and sharpen the rejected side through a taxonomy of regulatory failure modes : law swaps, article swaps, hallucinated citations, and hallucinated text. By fine-tuning on the resulting preference pairs exposes a divergence at the core of our findings: supervised fine-tuning attains the highest judge score of any method yet got the lowest rule-based Citation F1. SFT is producing longer answers with three times as many citations, most of them unsupported. DPO and GRPO offer the better trade-off with more concise answers and the highest citation F1 among fine-tuned models. Standard LLM-as-a-Judge evaluation rewards citation density as evidence of grounding and cannot detect when those citations are fabricated. Thus, we detect a blind spot that matters wherever answers must be verifiable. We release the benchmark and evaluation stack at https://github.com/auliakharis/FinRegQA-EU.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑