arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.23507cs.CLcs.AI

当名称跨越文字体系:面向蒙古世界历史实体对齐的来源基准

When Names Cross Scripts: A Source-Grounded Benchmark for Historical Entity Reconciliation in the Mongol World

Xiang Chen, Zeyu Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

针对蒙古世界历史实体对齐的MHER基准,证实来源基准证据可大幅提升对齐准确率,且仅名称输入在同名异义案例中表现极差,仅上下文信息与Qwen3-8B的名称处理存在特定问题。

中文摘要 AI 辅助

历史人物可能以不同语言、文字体系和转录传统出现,而不同个体可能拥有高度相似甚至完全相同的名称,这使得历史身份对齐问题远不止于字符串匹配或音译。我们推出MHER,这是一个用于蒙古世界人物名称证明成对对齐的来源控制基准。MHER包含一个由84位主要历史人物构成的、共396对的仅名称核心集,以及一个更严格的、基于逐来源提及证据构建的160对来源基准子集,且拥有实体不重叠的开发集和测试集。在五个生成式系统中,与仅名称输入相比,正确的来源基准证据使配对测试准确率提升了12.96至94.44个百分点。在25个表面形式相同但属于不同人物的案例中,所有模型在仅使用名称时均失败(0/25的模型-项决策),而来源基准证据则实现了24/25的正确分辨,剩余结果为弃权(不执行)。仅上下文的消融实验表明,历史描述通常携带大量身份信息,而明确标记的错误来源控制会显著降低性能。我们还发现名称并非始终有益:对于Qwen3-8B,恢复表面形式会将原本由仅上下文实现的10次正确区分转为错误的身份合并。这些结果表明,历史实体对齐不仅依赖表面对应关系,还取决于身份判断是否恰当响应来源控制的历史证据。因此,MHER为研究历史自然语言处理中的证据使用、弃权(不执行)及失败模式提供了受控框架。

英文摘要

Historical people may appear under different languages, scripts, and transcription traditions, while distinct individuals may share highly similar or even identical names. This makes historical identity reconciliation more than a problem of string matching or transliteration. We introduce MHER, a provenance-controlled benchmark for pairwise reconciliation of person-name attestations from the Mongol world. MHER contains a balanced 396-pair Name-only core over 84 primary historical persons and a stricter 160-pair Source-grounded subset constructed from mention-by-source evidence, with entity-disjoint development and test splits. Across five generative systems, correctly Source-grounded evidence improves paired TEST accuracy by 12.96 to 94.44 percentage points relative to Name-only input. On five identical-surface different-person cases, all models fail under names alone (0/25 model-item decisions), whereas Source-grounded evidence yields 24/25 correct resolutions, with the remaining output an abstention. Context-only ablations show that historical descriptions often carry substantial identity information, while explicitly signaled misgrounding controls produce substantially lower performance. We also find that names are not uniformly beneficial: for Qwen3-8B, restoring surface forms converts ten otherwise correct Context-only distinctions into false identity merges. These results show that historical entity reconciliation depends not only on surface correspondence, but on whether identity judgments respond appropriately to provenance-controlled historical evidence. MHER therefore provides a controlled framework for studying evidence use, abstention, and failure modes in historical NLP.

发表机构

  • University of Amsterdam(阿姆斯特丹大学)
  • Amsterdam UMC(阿姆斯特丹大学医学中心)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑