HalluTruthQA-4K:面向阿拉伯语幻觉检测与事实验证的细粒度语料库及标注流程
HalluTruthQA-4K: A Fine-Grained Corpus and Annotation Process for Arabic Hallucination Detection and Truth Verification
浏览论文内容
中文总结 AI 辅助
本文提出HalluTruthQA-4K,这一含4000个阿拉伯语问答实例的细粒度语料库,用于阿拉伯语幻觉检测等任务,为HalluScoring 2026共享任务提供官方数据集。
中文摘要 AI 辅助
大型语言模型能生成流畅的阿拉伯语回答,但会引入难以识别和验证的事实错误。现有的阿拉伯语幻觉资源通常对整个回答赋予二元标签,表明其是否存在幻觉,但几乎未提供确切错误内容、错误原因或正确事实答案的相关信息。本文提出HalluTruthQA-4K,它是HalluTruthQA资源的扩展版本,包含4000个由专家精心整理的阿拉伯语问答实例,覆盖伊斯兰知识、历史、科学和地理四个知识密集型领域。作为HalluScoring 2026共享任务第2轨的官方数据集,HalluTruthQA-4K将原语料库扩展至4000个实例,每个实例配对一个阿拉伯语问题、一个模型生成的回答、一个经验证的参考答案以及五个合理的干扰项。存在幻觉的回答还标注了字符级错误跨度、人工撰写的解释以及分层幻觉类型。该语料库包含1643个幻觉回答和2357个非幻觉回答,共1843个标注错误跨度。本文描述了资源构建与标注方法,包括问题选择、受控答案生成、候选构建、专家标注、独立验证、裁决和质量控制,还记录了标注指南、分类体系、数据格式、标注者间一致性及语料库统计数据。HalluTruthQA-4K为幻觉检测、跨度级错误定位、解释生成、事实验证以及阿拉伯语模型事实可靠性的更广泛评估提供了可复用资源。
英文摘要
Large language models can generate fluent Arabic answers while introducing factual errors that are difficult to identify and verify. Existing Arabic hallucination resources often assign a binary label to an entire response, indicating whether it is hallucinated or non-hallucinated, but provide limited information about the exact erroneous content, the reason for the error, or the correct factual answer. We present HalluTruthQA-4K, an expanded version of the HalluTruthQA resource containing 4,000 expert-curated Arabic question-answering instances across four knowledge-intensive domains: Islamic knowledge, history, science, and geography. Serving as the official dataset for Track 2 of the HalluScoring 2026 shared task, HalluTruthQA-4K extends our original corpus to 4,000 instances. Each instance pairs an Arabic question with a model-generated response, a verified reference answer, and five plausible distractors. Hallucinated responses are additionally annotated with character-level erroneous spans, human-written explanations, and hierarchical hallucination types. The corpus contains 1,643 hallucinated and 2,357 non-hallucinated responses, with 1,843 annotated erroneous spans. We describe the resource construction and annotation methodology, including question selection, controlled answer generation, candidate construction, expert annotation, independent verification, adjudication, and quality control. We also document the annotation guidelines, taxonomy, data format, inter-annotator agreement, and corpus statistics. HalluTruthQA-4K provides a reusable resource for hallucination detection, span-level error localization, explanation generation, factual verification, and the broader evaluation of factual reliability in Arabic language models.