SignRAG:统一的无 gloss 手语翻译检索增强框架
SignRAG: Unified Retrieval-Augmented Gloss-Free Sign Language Translation
浏览论文内容
中文总结 AI 辅助
针对仅解码器 LLM 难以适配无 gloss 手语翻译的问题,提出 SignRAG 框架,通过分层预训练、目标域检索增强及 RUG-RFT 实现 SOTA 性能,是首个在 CSL-Daily 所有指标上优于 gloss 监督方法的无 gloss 手语翻译方法。
中文摘要 AI 辅助
当代仅解码器的大语言模型(LLM)在众多领域展现出强大能力。然而,现有的无 gloss 手语翻译(SLT)预训练范式大多围绕传统的编码器-解码器预训练语言模型设计,这限制了它们在仅解码器 LLM 上的直接适用性。为解决这一局限,我们提出 SignRAG,这是一个结合分层预训练、目标域检索增强和检索感知强化微调的统一框架。分层预训练首先学习基于语言学的手语表示,随后将手语编码器与 LLM 联合对齐,缓解跨模态优化失衡。在下游适配阶段,SignRAG 用基于参数的微调补充目标域检索库,该库提供实例特定的翻译线索。为确保检索上下文被恰当使用,我们进一步引入检索效用引导的强化微调(RUG-RFT),它结合翻译质量和检索效用奖励,鼓励有益的检索使用,同时抑制有害依赖。在多个 SLT 基准上的实验取得了新的 SOTA 性能。尤其值得注意的是,据我们所知,SignRAG 是首个在 CSL-Daily 的所有报告指标上均优于 gloss 监督方法的无 gloss 方法。我们的代码已在 GitHub 发布,同时发布了不同规模的模型以支持未来学术研究。
英文摘要
Contemporary decoder-only large language models (LLMs) have demonstrated strong capabilities across a wide range of domains. However, existing pretraining paradigms for gloss-free sign language translation (SLT) are largely designed around conventional encoder-decoder pretrained language models, which limits their direct applicability to decoder-only LLMs. To address this limitation, we propose SignRAG, a unified framework combining hierarchical pretraining, target-domain retrieval augmentation, and retrieval-aware reinforcement fine-tuning. Hierarchical pretraining first learns linguistically grounded sign representations and then jointly aligns the sign encoder with an LLM, mitigating cross-modal optimization imbalance. For downstream adaptation, SignRAG complements parameter-based fine-tuning with a target-domain retrieval gallery that provides instance-specific translation cues. To ensure that retrieved contexts are used appropriately, we further introduce Retrieval Utility-Guided Reinforcement Fine-Tuning (RUG-RFT), which combines translation-quality and retrieval-utility rewards to encourage beneficial retrieval use while suppressing harmful reliance. Experiments on multiple SLT benchmarks establish new state-of-the-art performance. In particular, to the best of our knowledge, SignRAG is the first gloss-free approach to outperform gloss-supervised methods across all reported metrics on CSL-Daily. Our code has been released at \href{https://github.com/shahelaojieraozhi/SignRAG}{GitHub}, together with models of different sizes to support future academic research.
发表机构
- Macau University of Science and Technology(澳门科技大学)
- University of Macau(澳门大学)
- Chinese Academy of Sciences(中国科学院)
- Yale University(耶鲁大学)
- VIVO AI Lab(vivo AI实验室)
机构由 AI 辅助整理,请以论文原文为准。