arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从单文档到跨文档:大语言模型多粒度事件分析的基准测试

From Single- to Cross-Document: Benchmarking Multi-Granularity Event Analysis of Large Language Models

Tao Wen, Shuai Shao, Pei Ke, Xu Han, Jie Zou, Guannan Li, Tao Tian, Jinjie Qiu, Lan Wang, Ke Qin

arXiv 2607.27654首次发表:更新:

发表机构

University of Electronic Science and Technology of China; Tsinghua University(电子科技大学; 清华大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文推出MiGUE-Bench基准及MiGUE-Pipeline框架,通过四项核心任务评估LLMs多粒度事件分析能力,明确其能力边界与缺陷,为该领域改进提供方向。

AI 中文摘要

事件分析是信息抽取的重要基础方向,涉及不同文档粒度下的各类以事件为中心的任务。尽管大语言模型(LLMs)已在部分此类任务中初步取得良好性能,但受限于现有基准的文档粒度、任务设计及数据源,其在事件分析中的能力仍缺乏全面认知。为解决这些局限,本文推出MiGUE-Bench,这一用于评估LLMs多粒度事件分析性能的系统性基准。为支持大规模评估,我们首先开发名为MiGUE-Pipeline的LLM驱动自校正标注框架,实现高质量事件源数据及自动标签的可扩展获取;随后在基准中设计事件检测、关系推理、结构归纳、未来预测四项核心任务,以从原子事件细节到复杂跨文档叙事的不同层面探究模型能力。对最先进LLMs及检索增强生成(RAG)方法的大量实验,明确了当前能力边界并识别关键缺陷,为LLMs在高难度事件分析任务的未来改进提供洞见。

英文摘要

Event analysis is an essential and fundamental direction of information extraction, involving various event-centric tasks at different granularity of documents. While large language models (LLMs) have preliminarily achieved promising performance in part of these tasks individually, their capability in event analysis still lacks comprehensive understanding due to restricted document granularity, task designs, and data source of existing benchmarks. To address these limitations, we introduce MiGUE-Bench, a systematic benchmark for assessing the performance of LLMs in multi-granularity event analysis. To support large-scale evaluation, we first develop an LLM-driven self-correcting annotation framework called MiGUE-Pipeline, enabling scalable acquisition of high-quality source data of events with automatic labels. Then, we design four core tasks in our benchmark, i.e., event detection, relation reasoning, structure induction, and future prediction, to probe model competence at different levels, from atomic event details to complex cross-document narratives. Extensive experiments on state-of-the-art LLMs and retrieval-augmented generation (RAG) methods delineate the current capability boundary and identify critical deficiencies, providing insights into the future improvement of LLMs in challenging event analysis tasks.

Comments9 pages. Published in the Proceedings of the 49th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR 2026)

Journal refProceedings of the 49th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR '26), pp. 3464-3472, 2026

DOI:10.1145/3805712.3808607

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑