arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过添加少量SALT改进跨语言词元表示

Improving Cross-Lingual Token Representations by Adding a Pinch of SALT

Guillem Ramírez

arXiv 2609.09953首次发表:更新:

发表机构

ILCC, University of Edinburgh(爱丁堡大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出轻量级后训练方法SALT,通过跨度级监督改进跨语言句子编码器的词元表示,在多个多语言基准上取得最优结果,同时提升句子级性能。

AI 中文摘要

跨语言句子编码器能够实现跨数百种语言的可扩展迁移,为翻译挖掘和低资源环境下的零样本学习等应用提供支持。尽管这些编码器是为句子级对齐而训练的,但它们也越来越多地被应用于词元级任务,如幻觉检测和序列标注,这暴露了训练与使用之间的不匹配。我们提出了SALT,一种轻量级的后训练方法,通过向现有句子编码器中注入跨度级监督来改进词元表示。在五个多语言词元级基准测试中,SALT在四个基准上取得了最佳总体结果,优于其他微调策略和具有竞争力的编码器。它还在跨语言检索和分类任务上提升了句子级性能。这些结果表明,跨度级监督是改进词元表示和句子表示的有效信号。

英文摘要

Cross-lingual sentence encoders enable scalable transfer across hundreds of languages, powering applications such as translation mining and zero-shot learning in low-resource settings. Although trained for sentence-level alignment, they are increasingly also applied to token-level tasks such as hallucination detection and sequence tagging, exposing a mismatch between training and usage. We propose SALT, a lightweight post-training method that improves token representations by injecting span-level supervision into existing sentence encoders. Across five multilingual token-level benchmarks, SALT achieves the best overall results on four of them, outperforming alternative fine-tuning strategies and competitive encoders. It also improves sentence-level performance on cross-lingual retrieval and classification tasks. These results demonstrate that span-level supervision is an effective signal for improving both token and sentence representations.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑