MGAL:一个多语言粒度感知的长基准测试集
MGAL: A Multilingual Granularity-Aware Long-Context Benchmark
浏览论文内容
中文总结 AI 辅助
MGAL是首个多语言粒度位置感知长基准,以联合国6种官方语言报告构建,含4种粒度及位置分层,实验发现LLM粗粒度任务表现差、闭源模型低资源语言占优,还识别出语义拥挤等新挑战。
中文摘要 AI 辅助
长上下文大型语言模型(LLM)的评估已取得快速进展,但现有多数基准测试集局限于文档级别,且主要聚焦高资源语言,致使许多细粒度挑战未得到充分评估。为填补这一空白,我们提出MGAL,这是首个多语言、粒度及位置感知的长上下文基准测试集。MGAL由联合国(UN)报告构建,涵盖联合国6种官方语言,长度在8K至128K token之间,包含4个连贯的语言粒度级别(词、句子、段落、文档),并按条目在文档中的位置(开头、中间、结尾)进一步分层,同时在文档和段落级别建立索引。该设计可系统诊断不同粒度下的多语言长上下文理解能力。通过大量实验与分析,我们发现:(1)LLM在词级任务中表现良好,但在更粗粒度任务中表现不佳;(2)闭源模型在低资源语言中仍保持明显的性能优势。我们还识别出两个新挑战:(1)局部语义拥挤时,相邻句子共享主题与实体,模型倾向于遵循表面线索(如“however”等连接词或重复实体),而非句子在周围语境中的话语角色(如背景、结果);(2)生成输出的流畅性与一致性存在差距,模型生成的文本读起来流畅却偏离源事实。此外,我们观察到一些与先前研究一致的模式,包括依赖附近证据、不确定时复用选项。
英文摘要
Evaluation of long-context Large Language Models (LLMs) has advanced rapidly. However, most existing benchmarks are limited to the document level and focus mainly on high-resource languages, leaving many fine-grained challenges insufficiently evaluated. To address this gap, we present MGAL, the first multilingual, granularity- and position-aware long-context benchmark. MGAL is constructed from United Nations (UN) reports spanning 8K to 128K tokens across the six official UN languages. It covers four coherent levels of linguistic granularity (word, sentence, paragraph, and document) and further stratifies entries by their position within the document (begin, middle, and end), indexed at both the document and paragraph levels. This design enables systematic diagnosis of multilingual long-context comprehension across different granularities. Through extensive experiments and analyses, we find that: (1) LLMs perform well at word-level tasks but struggle with coarser-grained ones; and (2) Closed-source models retain a clear performance advantage in lower-resource languages. We further identify two new challenges: (1) Under local semantic crowding, where neighboring sentences share topics and entities, models tend to follow surface cues (e.g., connectives like ``however'' or repeated entities) rather than the discourse role of the sentence in surrounding context (e.g., background, outcome); and (2) A gap between fluency and consistency in generated outputs, where models produce text that reads smoothly but drifts from the source facts. In addition, we observe several patterns in line with prior studies, including reliance on nearby evidence and reuse of options under uncertainty.
发表机构
- China University of Petroleum (East China)(中国石油大学(华东))
- Northeastern University at Qinhuangdao(东北大学秦皇岛分校)
- Lingnan University, Hong Kong(香港岭南大学)
机构由 AI 辅助整理,请以论文原文为准。