HalluTruthQA:阿拉伯语问答中幻觉检测、定位和解释的细粒度基准
HalluTruthQA: A Fine-Grained Benchmark for Hallucination Detection, Localization, and Explanation in Arabic Question Answering
浏览论文内容
中文总结 AI 辅助
针对阿拉伯语问答中幻觉检测等难题,提出HalluTruthQA细粒度基准,含多领域2400个示例。在零样本设置下评估四个开源模型,发现各任务考查不同能力,强调幻觉评估应拓展到定位、验证和解释事实错误。
中文摘要 AI 辅助
大语言模型能生成流畅的阿拉伯语答案,但事实错误仍难检测、定位、解释和验证。现有幻觉基准通常提供响应级标签,对识别确切错误内容、解释错误原因或选择正确事实答案的支持有限。我们引入了HalluTruthQA,这是一个用于阿拉伯语问答中幻觉评估的细粒度基准。该基准包含2400个经过专家策划的示例,涵盖四个知识密集型领域。我们在零样本设置下对四个开源大语言模型进行评估,结果表明不同任务考查不同能力,没有单一模型在所有任务中都表现最强。我们的分类法表明,幻觉评估应从检测转向定位、验证和解释事实错误。代码、数据集等可通过链接获取。
英文摘要
Large language models (LLMs) can generate fluent Arabic answers, yet factual errors remain difficult to detect, localize, explain, and verify. Existing hallucination benchmarks often provide response-level labels, with limited support for identifying the exact erroneous content, explaining why it is incorrect, or selecting the correct factual answer. We introduce HalluTruthQA, a fine-grained benchmark for hallucination evaluation in Arabic question answering. The benchmark contains 2,400 expert-curated examples across four knowledge-intensive domains: Islamic knowledge, history, science, and geography. Each example pairs an Arabic question and a model-generated answer with a verified reference answer, a binary hallucination label, and six candidate answers for factual verification. Hallucinated answers additionally include character-level erroneous spans, human-written explanations, and macro- and micro-level hallucination types. We evaluate four open-source LLMs, ALLaM-7B, Falcon-H1R-7B, Qwen3-32B, and SILMA, in a zero-shot setting across hallucination detection, span-level localization, factual verification, and explanation evaluation. Results show that these tasks capture different abilities: no single model performs best across all tasks. The best scores are 0.880 Macro-F1 for detection, 0.516 F1-Sp for localization, 0.852 LO-Score for factual verification, and 0.644 for explanation evaluation. These findings show that hallucination evaluation should move beyond response-level detection toward the localization, verification, and explanation of factual errors.
发表机构
- Hamad Bin Khalifa University(哈马德·本·哈利法大学)
- Qatar University(卡塔尔大学)
- University of Biskra(比斯克拉大学)
- University of the Basque Country(巴斯克大学)
- Nazarbayev University(纳扎尔巴耶夫大学)
- Universiti Malaysia Kelantan(马来西亚吉兰丹大学)
机构由 AI 辅助整理,请以论文原文为准。