发表机构
State Key Laboratory of Novel Software Technology, Nanjing University; University of Warwick(南京大学计算机软件新技术国家重点实验室; 华威大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对无词素手语翻译(GFSLT)中LLMs的视觉特征利用问题,提出SignLlama模型,通过过滤伪词素CTC预训练与视觉优先蒸馏策略,在多数据集上取得极具竞争力的性能。
AI 中文摘要
大语言模型(LLMs)在各类任务中已取得显著成功,但针对无词素手语翻译(GFSLT)对LLMs进行微调仍是一项挑战。本文研究如何有效适配LLMs至GFSLT任务,明确需解决两个关键问题:(1)视觉特征输入与文本特征输入之间固有的分布差距,导致LLMs难以解释视觉输入;(2)现有方法通常在自回归框架中拼接视觉与文本特征,由于LLMs主要在以文本为中心的数据上预训练,模型会过度强调文本输入而轻视视觉线索。为解决第一个挑战,本文提出一种简单有效的方法——过滤伪词素连接时序分类预训练(Filtered Pseudo-Gloss CTC Pretraining),利用从文本序列生成的过滤伪词素序列监督视觉骨干网络的训练。为解决第二个问题,本文引入视觉优先蒸馏训练策略,具体而言,定义一条仅视觉的预测路径,其中文本输入被掩码,模型需仅依赖视觉输入生成目标序列;为引导该路径,将标准视觉-文本预测的输出蒸馏至仅视觉的预测路径,促使模型优先考虑视觉特征。综合实验与定性分析证明了所提模型的有效性,本文提出的SignLlama在多个GFSLT任务数据集上实现了极具竞争力的性能,且未使用任何额外模态或外部手语数据集进行预训练。
英文摘要
Large Language Models (LLMs) have achieved remarkable success across a wide range of tasks. However, fine-tuning LLMs for Gloss-Free Sign Language Translation (GFSLT) remains a challenge. In this paper, we investigate how to effectively adapt LLMs to the GFSLT task. We show that there are two key issues that need to be solved: (1) the inherent distributional gap between visual feature inputs and text feature inputs makes it difficult for LLMs to interpret visual inputs; and (2) existing approaches typically concatenate visual and textual features in an autoregressive framework, which leads to the model overemphasizing textual inputs and deprioritizing visual cues, as LLMs are pretrained predominantly on text-centric data. To address the first challenge, we propose a simple yet effective method named Filtered Pseudo-Gloss CTC Pretraining, which leverages filtered pseudo-gloss sequences generated from text sequences to supervise the training of the visual backbone. To tackle the second issue, we introduce a Visual-Prioritized Distillation training strategy. Specifically, we define a visual-only prediction path in which text inputs are masked, and the model is required to generate the target sequence relying solely on visual inputs. To guide this path, the outputs from the standard visual-textual prediction are then distilled into the visual-only prediction path, encouraging the model to prioritize visual features. Comprehensive experiments and qualitative analyses demonstrate the effectiveness of the proposed model. The proposed SignLlama achieves very competitive performance on multiple datasets for GFSLT tasks, without using any extra modalities or external sign language datasets for pretraining.