arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过结构相似性改进研究论文中序列句子分类的跨语言迁移

Improving Cross-Lingual Transfer for Sequential Sentence Classification in Research Papers via Structural Similarity

Kazuhiro Yamauchi, Marie Katsurai

arXiv 2609.19650首次发表:更新:

发表机构

Doshisha University(同志社大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过构建多语言数据集和实验,发现结构相似性比语言邻近性更能预测序列句子分类的跨语言迁移性能,并提出利用生成模型显式利用结构信息的方法,在域内和跨语言迁移中达到或超越最强基线。

AI 中文摘要

序列句子分类(SSC)是结构化科学文献的一项基本任务,将SSC研究扩展到英语以外的语言,可以提高多语言数字图书馆中科学知识的可获取性。跨语言迁移是解决非英语语言训练数据稀缺问题的一种有前景的方法。先前在其他自然语言处理任务上的工作表明,捕捉源语言和目标语言之间的语言相似性是有益的。然而,SSC本质上依赖于话语层面的模式,如标签序列和位置规律性,这些模式在不同语言中一致出现,不受语言差异的影响。为了考察决定SSC迁移成功的因素,我们构建了一个涵盖13种非英语语言的多语言SSC数据集,这些数据来自五个学术数据库。我们使用基于编码器的模型和生成模型进行的跨语言迁移实验表明,语言邻近性对迁移性能没有一致的预测能力,而修辞组织中的结构相似性则显示出弱但一致的积极相关性,且这种相关性在不同模型中均存在。在控制源语言性能后,标签分布的相似性是最一致的预测因子。基于这一发现,我们提出了一组三种方法,利用生成模型显式地利用结构信息。在域内评估中,最佳组合达到了与最强编码器基线相当的性能,而在对训练期间未见语言的迁移中,它优于最强的编码器基线。

英文摘要

Sequential sentence classification (SSC) is an essential task for structuring scientific publications, and extending SSC research to languages other than English can improve accessibility to scientific knowledge in multilingual digital libraries. Cross-lingual transfer is a promising approach to address the scarcity of training data in non-English languages. Prior work on other natural language processing tasks has shown the benefits of capturing linguistic similarity between source and target languages. However, SSC inherently depends on patterns at the discourse level, such as label sequences and positional regularities, which appear consistently across languages regardless of linguistic differences. To examine the factors that determine transfer success in SSC, we constructed a multilingual SSC dataset covering 13 non-English languages collected from five academic databases. Our cross-lingual transfer experiments, using both encoder-based and generative models, show that linguistic proximity has no consistent predictive power for transfer performance, whereas structural similarity in rhetorical organization shows a weak but consistent positive correlation across models. After controlling for source-language performance, the similarity of label distributions is the most consistent predictor. Building on this finding, we propose a set of three methods that explicitly leverage structural information using generative models. In the in-domain evaluation, the best combination reaches parity with the strongest encoder baselines, and in transfer to languages unseen during training, it outperforms the strongest encoder baseline.

CommentsAccepted at JCDL 2026 (ACM/IEEE Joint Conference on Digital Libraries), Frisco, TX, USA, October 13-16, 2026. 12 pages, 5 figures, 9 tables. DOI: 10.1145/3805696.3846040

DOI:10.1145/3805696.3846040

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑