发表机构
WPS Qingqiu(WPS清qiu(WPS清秋))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
OmniAlign是一款轻量级多语言对齐工具,采用四阶段训练流程,可同时支持词级和句子级对齐,在基准测试中表现具竞争力且泛化性良好,短文本微调可提升其对齐质量与长文本稳健性。
AI 中文摘要
跨语言序列对齐是构建和利用平行语料库的基础,涵盖从文档、句子到词和子词的映射。然而,现有工具通常仅针对单一粒度进行优化,因此从业者往往需要单独的系统来处理词级和句子级对齐——尤其在多语言和长文本场景中。我们提出OmniAlign,一款统一的多语言对齐工具,仅用一个轻量级模型即可同时支持词级和句子级对齐。OmniAlign基于具备强长文本建模能力的编码器架构构建,通过上下文词元相似度矩阵生成词对齐,结合句子嵌入与动态规划获得文档级的m-n句子对齐。为平衡细粒度对齐精度与句子表示质量,我们采用四阶段训练流程:面向对齐的持续预训练、自监督学习、基于人工标注的监督微调,以及从强大多语言教师模型进行的句子嵌入蒸馏。实验表明,OmniAlign在词级和句子级对齐基准上均取得极具竞争力的性能,且能很好地泛化到未见过的语言对。令人惊讶的是,在短文本上进行的后期监督微调进一步提升了对齐质量,同时保留了前期训练获得的长文本理解能力,使模型在长文本词对齐上保持稳健。\n代码:this https URL\n模型:this https URL
英文摘要
Cross-lingual sequence alignment is fundamental for building and exploiting parallel corpora, spanning mappings from documents and sentences down to words and subwords. Existing tools, however, typically specialize in a single granularity, so practitioners often need separate systems for word- and sentence-level alignment---especially in multilingual and long-text settings. We present OmniAlign, a unified multilingual aligner that supports both word-level and sentence-level alignment with a single lightweight model. Built on an encoder-only backbone with strong long-context modeling, OmniAlign induces word alignments from contextualized token similarity matrices, and obtains document-level $m$--$n$ sentence alignments via sentence embeddings combined with dynamic programming. To balance fine-grained alignment accuracy and sentence-representation quality, we use a four-stage training pipeline: alignment-oriented continued pre-training, self-supervised learning, supervised fine-tuning on human annotations, and sentence-embedding distillation from a strong multilingual teacher. Experiments show that OmniAlign achieves highly competitive performance on both word- and sentence-alignment benchmarks and generalizes well to unseen language pairs. Surprisingly, later-stage supervised fine-tuning on short texts further improves alignment quality while retaining the long-context understanding acquired in earlier training, keeping the model robust on long-text word alignment. \normalsize {\color{blue}\textbf{Code}: https://github.com/MilkDargon/OmniAlign}\par {\color{blue}\textbf{Model}: https://huggingface.co/WPS-Qingqiu/OmniAlign}