arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.17687cs.AIcs.LG

混合专家模块包含强幻觉检测信号

Mixture-of-Expert Blocks Contain Strong Hallucination Detection Signals

Joao Fonseca, Rodrigo Rodrigues, Paolo Romano

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出首个利用混合专家(MoE)专属内部信号的InnerExpert方法,在五个数据集和两种MoE架构上,以单次前向传播实现优于现有方法的逐词幻觉检测性能。

中文摘要 AI 辅助

尽管大语言模型(LLMs)已被广泛使用,但仍受限于一个根本问题:生成看似合理但实则错误的内容,即所谓的“幻觉”。现有大多数检测方法在答案或句子层面运行,然而逐词检测对于定位幻觉片段并实现细粒度干预至关重要。本文中,我们探索利用混合专家(Mixture-of-Experts, MoE)范式解决这一缺口。在MoE架构中,单次前向传播通过路由机制激活稀疏的专家子集(即每层中不同的前馈网络),产生密集架构中不存在的内部信号(如路由器熵、专家分歧和专家使用模式),这些信号此前未被用于幻觉检测。为此,我们提出InnerExpert,首个利用这些MoE专属信号进行逐词幻觉检测的方法。InnerExpert将路由级信号与标准Transformer信号结合为紧凑的逐词特征向量,由轻量级检测器分类,该检测器基于“LLM作为评判者”(LLM-as-a-judge)流程生成的标签训练,无需人工标注即可实现模型的连续更新。我们的结果显示,InnerExpert在五个数据集和两种MoE架构上的表现优于现有方法,在单次前向传播的情况下,达到了0.91的答案层面AUROC和0.76的词元层面AUROC。

英文摘要

Despite their widespread use, Large Language Models (LLMs) remain limited by a fundamental problem: the generation of plausible but false content, known as hallucinations. Most existing detection methods operate at the answer or sentence level, yet per-token detection is essential for localizing hallucinated spans and enabling fine-grained interventions. In this paper, we explore the use of the Mixture-of-Experts (MoE) paradigm to address this gap. In MoE architectures, a single forward pass activates a sparse subset of experts (i.e., distinct feedforward networks per layer) via a routing mechanism, producing internal signals (e.g., router entropy, expert disagreement, and expert usage patterns) that are unavailable in dense architectures and have not been previously exploited for hallucination detection. To this end, we introduce InnerExpert, the first method to leverage these MoE-specific signals for per-token hallucination detection. InnerExpert combines routing-level and standard transformer signals into compact per-token feature vectors, classified by a lightweight detector trained on labels produced by an LLM-as-a-judge pipeline, which enables continuous model updates without manual annotation. Our results show that InnerExpert outperforms existing methods across five datasets and two MoE architectures, achieving up to 0.91 answer-level and 0.76 token-level AUROC, while requiring only a single forward pass.

↑