发表机构
ILCC, University of Edinburgh(爱丁堡大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出轻量级后训练方法SALT,通过跨度级监督改进跨语言句子编码器的词元表示,在多个多语言基准上取得最优结果,同时提升句子级性能。
AI 中文摘要
跨语言句子编码器能够实现跨数百种语言的可扩展迁移,为翻译挖掘和低资源环境下的零样本学习等应用提供支持。尽管这些编码器是为句子级对齐而训练的,但它们也越来越多地被应用于词元级任务,如幻觉检测和序列标注,这暴露了训练与使用之间的不匹配。我们提出了SALT,一种轻量级的后训练方法,通过向现有句子编码器中注入跨度级监督来改进词元表示。在五个多语言词元级基准测试中,SALT在四个基准上取得了最佳总体结果,优于其他微调策略和具有竞争力的编码器。它还在跨语言检索和分类任务上提升了句子级性能。这些结果表明,跨度级监督是改进词元表示和句子表示的有效信号。
英文摘要
Cross-lingual sentence encoders enable scalable transfer across hundreds of languages, powering applications such as translation mining and zero-shot learning in low-resource settings. Although trained for sentence-level alignment, they are increasingly also applied to token-level tasks such as hallucination detection and sequence tagging, exposing a mismatch between training and usage. We propose SALT, a lightweight post-training method that improves token representations by injecting span-level supervision into existing sentence encoders. Across five multilingual token-level benchmarks, SALT achieves the best overall results on four of them, outperforming alternative fine-tuning strategies and competitive encoders. It also improves sentence-level performance on cross-lingual retrieval and classification tasks. These results demonstrate that span-level supervision is an effective signal for improving both token and sentence representations.