arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SFAD:推测性事实感知解码

SFAD: Speculative Factuality-Aware Decoding

Guanqiao Chen, Di Wang, Lijie Hu

arXiv 2609.00796首次发表:更新:

发表机构

MBZUAI; University of Science and Technology of China; King Abdullah University of Science and Technology(穆罕默德·本·扎耶德人工智能大学; 中国科学技术大学; 阿卜杜拉国王科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对大语言模型的上下文忠实性挑战,提出SFAD推测解码框架,构建ConFide数据集训练草稿模型,通过认知摩擦检测幻觉并优化分布,在提升忠实性的同时实现2.48倍加速。

AI 中文摘要

作为大语言模型最关键的挑战之一,上下文忠实性直接决定了其在知识密集型应用中的可靠性,该任务尤为棘手,因为需要在事实一致性与生成效率间取得平衡。对比解码方法需执行两次前向传播(带与不带上下文)以比较模型输出,使推理计算开销翻倍;而后训练对齐则需要大量计算开销的强化学习。为解决这一挑战,我们提出SFAD,一种提升上下文忠实性且不降低推理速度的推测解码框架。我们首先构建ConFide,一种带有细粒度原子扰动的偏好数据集,通过直接偏好优化训练上下文忠实的草稿模型。推理阶段,认知摩擦通过量化由专家确定性加权的分布张力检测潜在幻觉;当摩擦超过阈值时,非对称对数it引导通过基于残差的对数it注入优化目标分布,否则执行标准推测。大量实验表明,SFAD大幅提升忠实性,同时实现2.48倍加速,为高效大语言模型提供实用解决方案。

英文摘要

As one of the most critical challenges in large language models, contextual faithfulness directly determines their reliability in knowledge-intensive applications. This task is particularly challenging as it requires balancing factual consistency with generation efficiency. Contrastive decoding methods require dual forward passes (with and without context) to compare model outputs, doubling inference computational overhead, while post-training alignment demands extensive reinforcement learning with substantial computational overhead. To address this challenge, we present SFAD, a speculative decoding framework that enhances contextual faithfulness without inference degradation. We first construct ConFide, a preference dataset with fine-grained atomic perturbations, to train a context-faithful draft model via Direct Preference Optimization. During inference, Epistemic Friction detects potential hallucinations by quantifying distributional tension weighted by specialist certainty. When friction exceeds the threshold, Asymmetric Logit Steering refines the target distribution through residual-based logit injection; otherwise, standard speculation proceeds. Extensive experiments demonstrate that SFAD substantially improves faithfulness while achieving $2.48\times$ speedup, offering a practical solution for efficient LLMs.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑