arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

真实性信号在语码混合下是否幸存?探测隐藏状态以检测印地英语码混合中的幻觉

Does the Truthfulness Signal Survive Code-Mixing? Probing Hidden States for Hallucination Detection in Hinglish

Tanveer Singh

arXiv 2609.22138首次发表:更新:

发表机构

Plaksha University(普拉克沙大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究探究幻觉探测器在语码混合(Hinglish)下的迁移能力,发现信号幸存良好,且印地语训练探测器更可靠,并揭示模型在印地语和Hinglish上幻觉更多。

AI 中文摘要

隐藏状态幻觉探测——在LLM的内部激活上训练线性分类器,以检测生成的答案是否忠实于输入——是2026年研究的一个活跃领域,近期研究在多个基准和语言上报告了0.90-1.00的AUROC。然而,尽管大量聊天机器人用户使用印地语-英语混合文本(“Hinglish”)书写,这些研究均未在语码混合输入上测试探测器的性能。我们直接解决这一空白:在干净语言隐藏状态上训练的幻觉探测器能否迁移到Hinglish,还是信号在语码混合下会退化?我们构建了一个包含5,674个条目的印地语/英语/Hinglish问答基准,在三个开放权重7-8B LLM(Qwen2.5-7B、Mistral-7B、Llama-3.1-8B)上生成并标注了17,022个模型响应,提取两个token位置上的逐层隐藏状态,并训练线性和MLP探测器用于分布内检测和跨语言迁移。我们发现幻觉信号在语码混合下幸存良好:迁移AUROC范围从0.88到0.99,与分布内性能相比差距大多低于0.05 AUROC,并且印地语训练的探测器比英语训练的探测器更可靠地迁移到Hinglish。作为一个独立的、具有实际动机的发现,所有三个模型在印地语和Hinglish上的幻觉程度显著高于英语,对于匹配的事实。我们发布我们的代码和合成的Hinglish问答数据集,以支持语码混合幻觉检测的进一步工作。

英文摘要

Hidden-state hallucination probing - training a linear classifier on an LLM's internal activations to detect whether a generated answer is faithful to the input - is an active area of 2026 research, with recent work reporting 0.90-1.00 AUROC across several benchmarks and languages. However, none of this work has tested probes on code-mixed input, despite the fact that a huge population of chatbot users write in Hindi-English code-mixed text ("Hinglish"). We address this gap directly: does a hallucination probe trained on clean-language hidden states transfer to Hinglish, or does the signal degrade under code-mixing? We construct a 5,674-item Hindi/English/Hinglish QA benchmark, generate and label 17,022 model responses across three open-weight 7-8B LLMs (Qwen2.5-7B, Mistral-7B, Llama-3.1-8B), extract per-layer hidden states at two token positions, and train linear and MLP probes for in-distribution detection and cross-lingual transfer. We find that the hallucination signal survives code-mixing well: transfer AUROC ranges from 0.88 to 0.99, with gaps of mostly under 0.05 AUROC relative to in-distribution performance, and that Hindi-trained probes transfer to Hinglish more reliably than English-trained probes. As an independent, practically motivated finding, all three models hallucinate substantially more on Hindi and Hinglish than on English for matched facts. We release our code and synthetic Hinglish QA dataset to support further work on code-mixed hallucination detection.

Comments9 pages, 2 figures, 4 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑