arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

单次前向不确定性头用于波斯语医学语言模型中的声明级幻觉检测

Single-Pass Uncertainty Heads for Claim-Level Hallucination Detection in Persian Medical Language Models

Mehrdad Ghassabi, Pedram Rostami, Hamidreza Baradaran Kashani, Sadra Hakim, Audrina Ebrahimi

arXiv 2610.03482首次发表:更新:

发表机构

University of Isfahan; University of Tehran; University of Windsor; University of Texas at Dallas(伊斯法罕大学; 德黑兰大学; 温莎大学; 德克萨斯大学达拉斯分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究将LUH框架适配至波斯语医学模型,构建声明级数据集并训练轻量级头,实现单次前向的幻觉检测,PR-AUC达随机基线2.3倍以上。

AI 中文摘要

幻觉检测对于医学语言模型尤为重要,但重复采样方法成本高昂,且现有的不确定性头资源无法直接迁移到新的骨干模型和语言上。我们将LLM不确定性头(LUH)框架适配到基于Aya-Expanse-8B的波斯语医学模型中,使用Gaokerena-V和Gaokerena-R作为两个先前开发的骨干模型。我们首先在168道伊朗医学入学考试题目上检验响应变异性,观察到Gaokerena-V的五次运行一致性显著低于Aya-Expanse-8B,而Gaokerena-R与Aya-Expanse-8B相当。随后,我们直接构建了两个配对的声明级幻觉数据集(波斯语),每个骨干模型包含1,600个响应,并在冻结的骨干注意力图和词元概率上训练轻量级声明级头。在保留测试集上,这些头获得的PR-AUC分别为0.4820和0.4652,对应其各自随机基线的2.30倍和2.66倍,ROC-AUC分别为0.7852和0.7810。这些头在推理时既不需要检索,也不需要重复采样。这些结果为波斯语医学语言模型的单次前向声明级不确定性估计提供了初步研究;测试集规模较小,标签为自动生成。

英文摘要

Hallucination detection is particularly important for medical language models, but repeated-sampling approaches are computationally expensive. A faster alternative is a single-pass uncertainty head that predicts hallucination risk from a frozen generator's internal signals. Existing uncertainty heads consume backbone-specific features and tokenization, motivating adaptation when the backbone or language changes. We study two Persian medical models, Gaokerena-V and Gaokerena-R, on a 168-question Iranian medical entrance examination. Across five generations per question, Gaokerena-V produces the same option on only 14 questions and Gaokerena-R on 37, compared with 168 for Med-Gemma, indicating substantial response variability in the Gaokerena models. We therefore adapt the LLM Uncertainty Head (LUH) framework to these models and construct two paired claim-level hallucination datasets directly in Persian, with 1,600 responses per backbone. Lightweight heads are trained on frozen-backbone attention maps and token probabilities. On held-out test splits, the heads achieve PR-AUCs of 0.4820 and 0.4652, corresponding to 2.30 and 2.66 times their respective random baselines, and ROC-AUCs of 0.7852 and 0.7810. The resulting detectors require neither retrieval nor repeated sampling at inference.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑