发表机构
King Fahd University of Petroleum and Minerals; Imam Abdulrahman bin Faisal University(法赫德国王石油与矿产大学; 伊玛目阿卜杜勒拉赫曼·本·费萨尔大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究大语言模型内部幻觉检测信号能否跨语言和领域泛化,通过CrossHallu对六个模型的内部表征在生成式问答任务上进行评估,结果揭示了模型内部状态幻觉信号跨语言和领域转移的情况。
AI 中文摘要
近期大语言模型中的幻觉检测技术专注于从模型内部表征直接提取特征并训练分类器,多数内部状态幻觉检测技术主要在英语中探索。为此提出CrossHallu,在生成式问答任务上用六个模型的内部表征评估幻觉检测的跨语言和跨领域泛化,进行多种评估,揭示了相关情况,代码公开。
英文摘要
Recent hallucination detection techniques in large language models (LLMs) focus on directly extracting features from a model's internal representations and training a classifier on these features to detect hallucinations, demonstrating promising results. Notwithstanding this advancement, most internal-state hallucination detection techniques have been explored predominantly in English, raising the question of whether such internal signals generalize across different languages and domains. To address this gap, we present CrossHallu, the first study to evaluate the cross-lingual and cross-domain generalization of hallucination detection using internal representations from six LLMs on the generative question-answering task. We conduct a systematic Arabic <-> English evaluation using TruthfulQA, an Arabic translated version of TruthfulQA, and HalluScore. This evaluation encompasses monolingual training and testing, cross-lingual transfer, cross-domain transfer, and combined cross-lingual and cross-domain transfer. The results reveal that internal-state hallucination signals in LLMs transfer across languages and domains for most models, with cross-lingual performance highly dependent on both class separability and language alignment in the feature space, whereas cross-domain transfer within Arabic varies depending on the training and testing datasets used for the hallucination detector. The code is publicly available at https://github.com/aishaalansari57/CrossHal.