三思而后行:利用内部归因信号进行事实性解码
Look Before You Leap: Factual Decoding with Internal Attribution Signals
浏览论文内容
中文总结 AI 辅助
针对LLM幻觉雪球问题,提出利用内部归因信号的DescaPE解码框架,通过探针近似事实显著信号并调整候选评分,在五个基准上提升事实性且仅增加1.10倍延迟。
中文摘要 AI 辅助
幻觉仍然是大型语言模型(LLMs)中的一个关键挑战,其中早期的事实错误在自回归生成过程中通过雪球效应不断累积,而事后纠正或权重层面的干预都无法有效预防。我们提出了DescaPE(解码信号控制对抗路径错误雪球),这是一种解码框架,利用模型内部信号在推理时抑制易产生幻觉的轨迹。通过滑动窗口MLP消融,我们在LLMs中识别出一个事实显著性层区间,其派生信号对事实性标记选择性增强,并在易产生幻觉的步骤表现出异常尖峰。我们训练了一个轻量级探针,从单次前向传播中近似该信号,并将其集成到候选评分中,以惩罚高风险续写并奖励基于事实的续写。在三个LLM上的五个事实性基准上的实验表明,DescaPE在多种设置下相对于解码时基线实现了事实性改进,同时在我们的效率评估中仅产生1.10倍的延迟开销。我们的代码可在该https URL获取。
英文摘要
Hallucination remains a critical challenge in large language models (LLMs), where early factual errors compound through autoregressive generation in a snowballing effect that neither post-hoc correction nor weight-level intervention can effectively preempt. We propose DescaPE (DEcoding Signal Control Against Path Error-snowballing), a decoding framework that leverages internal model signals to suppress hallucination-prone trajectories at inference time. Through sliding-window MLP ablation, we identify a factual-salient layer span within LLMs whose derived signal is selectively elevated for factual tokens and exhibits anomalous spikes at hallucination-prone steps. We train a lightweight probe to approximate this signal from a single forward pass and integrate it into candidate scoring to penalize high-risk continuations while rewarding factually grounded ones. Experiments across five factuality benchmarks on three LLMs demonstrate that DescaPE achieves factuality improvements over decoding-time baselines in multiple settings, while incurring only 1.10x latency overhead in our efficiency evaluation. Our code is available at https://github.com/hayeonggg/DESCAPE.
发表机构
- Chung-Ang University(中央大学)
机构由 AI 辅助整理,请以论文原文为准。