发表机构
College of Computer Science and Technology, National University of Defense Technology(国防科技大学计算机科学与技术学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究探索混合专家模型中基于专家感知的对比解码以减轻大语言模型幻觉问题。提出EAACD方法,利用MoE高层专家差异,将专家分组对比校准预测,放大低可靠性专家幻觉提供负面参考,在四个数据集上优于所有基线。
AI 中文摘要
现有的大语言模型幻觉缓解方法,如提示工程和模型优化,要么几乎不改变模型内部知识,要么跨域泛化性差。对比解码通过利用大语言模型中的分层差异来减轻幻觉,但先前研究仅探索基于Transformer的模型,忽略了混合专家(MoE)模型等其他有效框架。本文进行实证研究,发现共享专家的MoE中不存在类似分层差异,但不同MoE的较高层在事实与非事实输出间存在不同专家激活模式。在此基础上提出EAACD,即基于专家感知的自适应对比解码,利用MoE较高层的专家差异减轻问答任务中的幻觉。该方法将高层专家按置信度和一致性分为高可靠性组和低可靠性组,通过对比高可靠性组与低可靠性组的预测来校准模型原始预测,并通过注意力和掩码放大低可靠性专家的幻觉以提供更强负面参考。实验表明EAACD在四个数据集上优于所有基线。
英文摘要
Existing LLM hallucination mitigation methods, including prompt engineering and model optimization, either hardly alter models'internal knowledge or have poor cross-domain generalization. Contrastive decoding mitigates hallucinations by using layer-wise differences in LLMs. However, prior studies only explore transformer-based models (e.g., GPT), ignoring other effective frameworks like mixture-of-experts (MoE) models. Since MoE alters the traditional transformer architecture, we conduct empirical studies to investigate whether similar layer-wise differences exist in MoEs. Our results show that they do not exist in MoE with shared experts; nevertheless, across different MoEs, higher layers exhibit distinct expert activation patterns between factual and non-factual outputs. Building on these, we propose EAACD, an expert-aware adaptive contrast decoding that uses expert differences in MoE's higher layers to mitigate hallucinations on QA tasks. EAACD splits high-layer experts into a higher-reliability group and several lower-reliability groups based on their confidence and consistency. It contrasts the higher-reliability group's prediction with each lower-reliability group's prediction to calibrate the model's original predictions. To strengthen this contrast, EAACD amplifies hallucinations from lower-reliability experts via attention and masking to provide stronger negative references. EAACD outperforms all baselines on four datasets.
CommentsAccepted by ACL2