XHotpotQA:多跳问答中跨语言知识组合的基准
XHotpotQA: A Benchmark for Cross-Lingual Knowledge Composition in Multi-Hop Question Answering
AI总结:
XHotpotQA是针对多跳问答跨语言知识组合的基准,通过构建带语言分配的证据依赖图,揭示多语言问答中语言边界错误,为相关系统提供诊断与评估平台。
AI中文摘要:
基于知识的多跳问答要求系统选择证据并组合相关事实,然而多语言基准通常会将整个示例翻译成单一语言,这会掩盖推理链中语言边界处的错误。我们引入XHotpotQA,这是一个针对混合语言证据的跨语言知识组合的受控基准。每个实例被建模为证据依赖图,其问题、桥接证据、承载答案的证据以及干扰项都有明确的语言分配。该经审核的资源包含15661个训练实例和7405个验证实例,具备句子级支持监督并提供干扰项。在验证集中,99.81%的条目存在问题到黄金证据的语言界面交叉,95.60%使用不同语言的黄金段落。在三个阅读器模型上,完全问题-证据不匹配对应的Unicode感知答案F1值比部分对齐低10.25至15.79个点,不同脚本证据对应的F1值低11.98至23.70个点;对应的自适应选择器的差异为1.71和1.78个点。因此,在这种提供候选的设计下,被评估的阅读器表现出与条件相关的缺陷远大于选择器。XHotpotQA为必须跨语言整合证据的基于知识的系统提供角色感知诊断、模块化评估和经审核的测试平台。
英文摘要:
Knowledge-intensive multi-hop question answering requires systems to select evidence and compose dependent facts, yet multilingual benchmarks usually translate an entire example into one language. This hides failures at language boundaries inside the reasoning chain. We introduce XHotpotQA, a controlled benchmark for cross-lingual knowledge composition over mixed-language evidence. Each instance is modeled as an evidence-dependency graph whose question, bridge evidence, answer-bearing evidence, and distractors have explicit language assignments. The audited resource contains 15,661 training and 7,405 validation instances, with sentence-level support supervision and supplied distractors. In validation, 99.81% of items cross the question-to-gold-evidence language interface and 95.60% use gold paragraphs in different languages. Across three reader artifacts, full question-evidence mismatch is associated with 10.25 to 15.79 lower Unicode-aware answer F1 than partial alignment, and different-script evidence with deficits of 11.98 to 23.70 points; the corresponding adapted-selector contrasts are 1.71 and 1.78 points. Under this supplied-candidate design, the evaluated readers therefore show substantially larger condition-associated deficits than the selector. XHotpotQA provides role-aware diagnostics, modular evaluation, and an audited test bed for knowledge-based systems that must integrate evidence across languages.