arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.11371cs.CLcs.CVcs.MM

SignRAG:统一的无 gloss 手语翻译检索增强框架

SignRAG: Unified Retrieval-Augmented Gloss-Free Sign Language Translation

Zhi Rao, Yucheng Zhou, Qianran Sun, Yiqing Huang, Longcan Yuan, Jiayi Hou, Chengwen Yao, Lin Cheng, Donghui Sun, Xiaoxin Chen, Jun Wan

首次发表
浏览论文内容

中文总结 AI 辅助

针对仅解码器 LLM 难以适配无 gloss 手语翻译的问题,提出 SignRAG 框架,通过分层预训练、目标域检索增强及 RUG-RFT 实现 SOTA 性能,是首个在 CSL-Daily 所有指标上优于 gloss 监督方法的无 gloss 手语翻译方法。

中文摘要 AI 辅助

当代仅解码器的大语言模型(LLM)在众多领域展现出强大能力。然而,现有的无 gloss 手语翻译(SLT)预训练范式大多围绕传统的编码器-解码器预训练语言模型设计,这限制了它们在仅解码器 LLM 上的直接适用性。为解决这一局限,我们提出 SignRAG,这是一个结合分层预训练、目标域检索增强和检索感知强化微调的统一框架。分层预训练首先学习基于语言学的手语表示,随后将手语编码器与 LLM 联合对齐,缓解跨模态优化失衡。在下游适配阶段,SignRAG 用基于参数的微调补充目标域检索库,该库提供实例特定的翻译线索。为确保检索上下文被恰当使用,我们进一步引入检索效用引导的强化微调(RUG-RFT),它结合翻译质量和检索效用奖励,鼓励有益的检索使用,同时抑制有害依赖。在多个 SLT 基准上的实验取得了新的 SOTA 性能。尤其值得注意的是,据我们所知,SignRAG 是首个在 CSL-Daily 的所有报告指标上均优于 gloss 监督方法的无 gloss 方法。我们的代码已在 GitHub 发布,同时发布了不同规模的模型以支持未来学术研究。

英文摘要

Contemporary decoder-only large language models (LLMs) have demonstrated strong capabilities across a wide range of domains. However, existing pretraining paradigms for gloss-free sign language translation (SLT) are largely designed around conventional encoder-decoder pretrained language models, which limits their direct applicability to decoder-only LLMs. To address this limitation, we propose SignRAG, a unified framework combining hierarchical pretraining, target-domain retrieval augmentation, and retrieval-aware reinforcement fine-tuning. Hierarchical pretraining first learns linguistically grounded sign representations and then jointly aligns the sign encoder with an LLM, mitigating cross-modal optimization imbalance. For downstream adaptation, SignRAG complements parameter-based fine-tuning with a target-domain retrieval gallery that provides instance-specific translation cues. To ensure that retrieved contexts are used appropriately, we further introduce Retrieval Utility-Guided Reinforcement Fine-Tuning (RUG-RFT), which combines translation-quality and retrieval-utility rewards to encourage beneficial retrieval use while suppressing harmful reliance. Experiments on multiple SLT benchmarks establish new state-of-the-art performance. In particular, to the best of our knowledge, SignRAG is the first gloss-free approach to outperform gloss-supervised methods across all reported metrics on CSL-Daily. Our code has been released at \href{https://github.com/shahelaojieraozhi/SignRAG}{GitHub}, together with models of different sizes to support future academic research.

发表机构

  • Macau University of Science and Technology(澳门科技大学)
  • University of Macau(澳门大学)
  • Chinese Academy of Sciences(中国科学院)
  • Yale University(耶鲁大学)
  • VIVO AI Lab(vivo AI实验室)

机构由 AI 辅助整理,请以论文原文为准。

↑