arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DualAnchor:无注释手语翻译中保留语言先验并提升词汇保真度

DualAnchor: Preserving Language Priors and Improving Lexical Fidelity in Gloss-Free Sign Language Translation

Hongbin Zhang, Junhao Liu, Xuefeng Bai, Youcheng Pan, Yang Xiang, Kehai Chen

arXiv 2607.27614首次发表:更新:

AI 中文总结

针对现有无注释手语翻译方法的语言先验退化与词汇保真度差距问题,提出结合TPA和OTA的DualAnchor框架,在两个基准数据集上取得良好性能。

AI 中文摘要

大型语言模型(LLM)的最新进展使手语翻译(SLT,即将手语视频转换为口语文本的任务)越来越多地采用LLM作为文本主干。然而,尽管现有基于LLM的SLT方法具备强大的语言建模能力,却往往破坏而非利用这种语言先验,生成不流畅的译文,我们将这种失败称为语言先验退化。同时,现有方法通常在句子层面对齐视频与文本,这无法确保准确的词汇细节,造成了词汇保真度差距。为解决这两个问题,我们提出DualAnchor,这是一种无注释、基于LLM的SLT训练框架,结合两个互补锚点以实现语言流畅且视觉忠实的生成。令牌级先验锚定(TPA)通过在每个解码步骤将多模态解码器正则化为相同自回归前缀条件下冻结LLM的下一个令牌分布,从而保留LLM的语言先验。最优传输对齐(OTA)通过将视觉-文本匹配表述为熵正则化的部分最优传输,利用Sinkhorn优化在余弦代价下诱导视觉令牌与文本内容令牌之间的软对齐,以此提升词汇保真度。DualAnchor在PHOENIX-2014T和CSL-Daily上均实现了强劲的整体性能。针对性分析表明,这些提升源于两个锚点的互补效应:TPA提升流畅度,OTA减少细粒度词汇错误。

英文摘要

Recent advances in large language models (LLMs) have led sign language translation (SLT), the task of converting sign-language videos into spoken-language text, to increasingly adopt LLMs as textual backbones. However, despite their strong language modeling capabilities, existing LLM-based SLT methods often undermine rather than exploit this language prior, producing disfluent translations, a failure we term language-prior degradation. Meanwhile, existing methods typically align videos and text at the sentence level, which does not ensure accurate lexical details and creates a lexical fidelity gap. To address both issues, we propose DualAnchor, a gloss-free LLM-based SLT training framework that couples two complementary anchors for linguistically fluent and visually faithful generation. Token-level Prior Anchoring (TPA) preserves the LLM's language prior by regularizing the multimodal decoder at each decoding step toward the next-token distribution of a frozen LLM conditioned on the same autoregressive prefix. Optimal Transport Alignment (OTA) improves lexical fidelity by formulating visual-textual matching as entropy-regularized partial optimal transport, with Sinkhorn optimization inducing a soft alignment between visual tokens and textual content tokens under a cosine cost. DualAnchor achieves strong overall performance on both PHOENIX-2014T and CSL-Daily. Targeted analyses attribute these gains to the complementary effects of the two anchors: TPA improves fluency, whereas OTA reduces fine-grained lexical errors.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑