arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LongNovel:面向长语境小说摘要中幻觉检测的多尺度基准

LongNovel: A Multi-Scale Benchmark for Hallucination Detection in Long-Context Novel Summarization

Ruizhi Zhang, Jinwei Chen, Xiangju Lu, He Yan, Mo Yu, Junmin Zhu, Wei Zhang

arXiv 2608.18082首次发表:更新:

发表机构

East China Normal University; iQIYI Inc; Tencent(华东师范大学; 爱奇艺公司; 腾讯)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对长语境小说摘要幻觉检测缺乏多尺度基准的问题,构建了LongNovel双语多尺度基准,经实验验证其具有挑战性并将发布供后续研究使用。

AI 中文摘要

尽管近年来上下文窗口已显著扩大,但长语境摘要中的幻觉问题仍然是一项挑战。与新闻或论文相比,长篇小说更适合研究这些幻觉,因为其包含内在信息以及对事件和对话的详细描述。然而,当前研究缺乏面向长语境小说摘要中幻觉检测的多尺度基准,也未充分探究幻觉如何随上下文变长而变化。本研究提出LongNovel,这是一个面向幻觉检测的多尺度长语境双语(中英)小说基准,由29部中文小说(范围为16k至100k tokens)和BookSum数据集的章节级数据构建而成。我们设计了8种幻觉类型,并采用多模型仲裁(Multi-Model Arbitration)与实体参考幻觉生成(Entity-Referenced Hallucination Generation)相结合的方法,以确保数据真实性和幻觉类别的均衡分布;此外,我们手动修订测试集内容以保证数据可靠性。大量实验结果表明,LongNovel是一项具有挑战性的基准,我们将发布LongNovel以供未来研究使用。

英文摘要

Although context windows have expanded significantly in recent years, hallucinations in long-context summarization remain a challenge. Long novels are better suited than news or papers for researching these hallucinations, due to their intrinsic information and detailed descriptions of events and dialogues. However, current research lacks a multi-scale benchmark for hallucination detection in long-context novel summarization and does not fully explore how hallucinations change as the context grows longer. In this study, we propose LongNovel, a multi-scale long-context bilingual (Chinese and English) novel benchmark for hallucination detection. This benchmark is constructed from 29 Chinese novels (ranging from 16k to 100k tokens) and chapter-level data from the BookSum dataset. We design 8 hallucination types and employ a combination of Multi-Model Arbitration and Entity-Referenced Hallucination Generation to ensure both data authenticity and a balanced distribution of hallucination categories. Furthermore, we manually revise the content in the test set to guarantee data reliability. Extensive experimental results demonstrate that LongNovel is a challenging benchmark. We release LongNovel for future research. https://github.com/BDML-lab/LongNovel

CommentsAccepted at EMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑