arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LODESTAR:可信赖熵是被引导的,而非仅被测量——强化偏振器防止冻结型大语言模型(LLM)被错误证据误导而自信出错

LODESTAR: Robust Entropy-Based Answer Selection in Retrieval-Augmented Generation for Question Answering -- Directing Frozen-LLM Entropy with a Reinforcement-Learned Prompt Polarizer under Misleading Passages

Hung-Chun Hsu, Po-Jen Ko, Che-Cheng Wu, Li-Yang Chang, Chuan-Ju Wang

arXiv 2608.11922首次发表:更新:

AI 中文总结

LODESTAR是首个通过强化学习训练偏振器、依据第三方冻结型LLM的不确定性评分文本干预的方法,可提升检索增强问答的F1值等指标,减少模型读取误导性段落的概率。

AI 中文摘要

在检索增强问答任务中,预测分布熵是一项强有力的选择规则:在五个问答基准测试中,保留冻结型应答大语言模型(LLM)生成的具有最低答案标记熵的候选答案,可在无标准答案的情况下,将平均答案F1值从0.4769提升至0.5148,且该结果优于检索器排名最高的段落。然而,先前基于熵的选择器所采用的这一最低熵规则,存在一种特定且关键的缺陷:误导性段落会使应答模型自信地给出错误答案,恰恰在看似最可信的信号处压低了其熵值。我们表明,该缺陷源于应答模型所读取的段落,而该段落的读取语境是我们可干预的输入。我们提出LODESTAR,据我们所知,这是首个通过对比同一问题的候选答案,依据第三方冻结型应答模型所产生的不确定性来对文本干预进行评分的方法。该方法使用强化学习离线训练一次偏振器——这是一个插入到应答模型提示中、永不进入其权重的简短固定自然语言字符串;其训练标签由离线的标准答案和两个LLM评判器构建,推理过程中不使用任何标签。在5008个问题上,针对相同的冻结型应答模型和相同的候选池评估所有竞争选择器,LODESTAR在所有可用于推理的选择器中达到最高平均F1值(0.5148至0.5339)、最高精确匹配率(0.4136)以及所评判的冻结型应答模型配置中最高的GPT-4o评判器分数(0.6435);其三次随机种子的平均F1值在70种方法-数据集组合中,相对于14种已发表配置均胜出,且每一种组合均通过配对检验具有显著性。该提升在域内和域外均成立,对偏振器进行消融实验表明,它能使应答模型更少读取误导性段落(26.0%对比30.3%)。

英文摘要

Predictive-distribution entropy is a strong answer-selection rule in retrieval-augmented generation (RAG) for question answering: across five QA benchmarks, selecting the answer a frozen respondent LLM produces with the lowest answer-token entropy lifts mean $F_1$ from 0.4769 to 0.5148 over the retriever's top-ranked passage, without gold answers. Yet this rule, which prior entropy-based selectors adopt, fails: a misleading passage makes the respondent confidently wrong, driving entropy down where the uncertainty signal looks most trustworthy. The failure comes from the passage the respondent reads, and the context it is read in is an input we can intervene on. We introduce LODESTAR: to our knowledge the first method to score a text intervention by the uncertainty it induces in a third-party frozen respondent, compared within one question. LODESTAR uses reinforcement learning (GRPO) to train, once and offline, a polarizer -- a short fixed natural-language string inserted into the respondent's prompt and never into its weights, directing entropy so that entropy-based answer selection stays robust to misleading passages; training labels are built from gold answers and two LLM judges, and inference reads neither. With every competing selector under the same frozen respondent and candidate pools on 5,008 questions, LODESTAR attains the highest mean $F_1$ of any inference-ready selector (0.5339), the highest macro exact match (0.4136), and the highest GPT-4o judge score of the frozen-respondent configurations judged (0.6435); its three-seed mean wins all 70 $F_1$ cells against fourteen published configurations and is paired-significant on $F_1$ against every one. The gain holds in-domain on NQ-Open and out-of-domain over SQuAD, TriviaQA, EntityQuestions and WebQuestions. Ablating the polarizer shows it is what makes the respondent read a misleading passage less often (26.0% vs 30.3%).

Comments28 pages, 3 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑