arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于主张提取与对齐的跨语言传记丰富

Cross-lingual Biography Enrichment via Claim Extraction and Alignment

Yifei Song, Ziyang Chen, Emil Sayilov, Claire Gardent

arXiv 2608.23390首次发表:更新:

发表机构

CNRS; LORIA; Université de Lorraine; Université Paris Dauphine - PSL; Université de Strasbourg(法国国家科学研究中心; 洛里亚实验室; 洛林大学; 巴黎多芬大学 - PSL大学; 斯特拉斯堡大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究推出CLAW-4L基准与基于主张的丰富框架,利用非英文维基百科传记的主张对齐证据丰富英文传记,发现其可提升英文传记覆盖度,但低资源场景仍有挑战。

AI 中文摘要

英文维基百科常被视为默认的百科来源,但非英文维基百科版本可为长尾人物提供更丰富的本地化信息。我们研究跨语言传记丰富任务:用同一人物的非英文传记所支持的事实,丰富现有英文传记。针对非英语语境下的女性人物,我们推出CLAW-4L基准,包含300个维基百科传记对,将英文传记与其法语、中文或阿塞拜疆语对应传记关联,同时配有主张标注和细粒度主张对关系语料库。我们提出基于主张的丰富框架,从两种传记中提取英文主张,对齐以识别非英文传记中的丰富证据,并利用选定主张重写英文传记。结果表明,非英文维基百科传记为提升英文传记覆盖度提供了有价值的证据,但低资源场景仍具挑战性。

英文摘要

English Wikipedia is often treated as the default encyclopedic source, yet non-English Wikipedia editions can contain richer locally grounded information for long-tail figures. We study cross-lingual biography enrichment: enriching an existing English biography with facts supported by a non-English biography about the same person. Focusing on women from non-English-speaking contexts, we introduce \textsc{CLAW-4L}, a benchmark consisting of 300 Wikipedia biography pairs linking an English biography with its French, Chinese or Azerbaijani counterpart, along with claim annotations and a fine-grained claim-pair relation corpus. We propose a claim-based enrichment framework that extracts English claims from both biographies, aligns them to identify enrichment evidence from the non-English biography, and rewrites the English biography using the selected claims. Our results show that non-English Wikipedia biographies provide valuable evidence for improving English biography coverage, while lower-resource settings remain challenging.

CommentsAccepted by EMNLP 2026 main conference

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑