arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.00235cs.CV

用于手语翻译的注意力引导视觉-语言模型

Attention-Steered Vision-Language Models for Sign Language Translation

Meibo Hu, Guohao Sun, Annemarie D. Ross, Sheng Li, Zhiqiang Tao

首次发表
浏览论文内容

中文总结 AI 辅助

针对视觉-语言模型在手语翻译中时空视觉定位差的问题,提出AttnSign框架,通过空间注意力监督和RL运动节奏引导提升性能,在How2Sign和OpenASL基准上表现优于现有方法。

中文摘要 AI 辅助

视觉-语言模型(VLMs)已成为多模态视频理解的强大框架,但在手语翻译任务中仍存在局限,我们发现现有基于VLM的翻译器存在一个关键失效模式:时空视觉定位能力差。具体而言,标准的下一个标记交叉熵无法直接为模型应在何处、何时关注提供信号,导致模型忽略与手语相关的区域和帧。为应对这一挑战,我们提出AttnSign,这是一种用于手语翻译的基于VLM的时空注意力引导框架。AttnSign首先为每一帧中与手语相关的区域(如面部和手部)引入空间注意力监督;随后开发了一种基于强化学习(RL)的运动节奏引导方法,鼓励模型探索并聚焦于手语级关键帧。在How2Sign和OpenASL基准上的实验表明,我们提出的AttnSign始终优于现有方法。

英文摘要

Vision-language models (VLMs) have emerged as a powerful framework for multimodal video understanding. However, they remain limited in the sign language translation task, where we identify a key failure mode of existing VLMbased translators: poor spatial-temporal visual grounding. In particular, we find that standard next-token cross-entropy does not directly provide signal for where and when the model should attend, causing models to overlook sign-relevant regions and frames. To address this challenge, we propose AttnSign, a VLM-based spatial-temporal attention steering framework for sign language translation. AttnSign first introduces spatial attention supervision for sign-relevant regions, such as face and hands, in each frame; then develops an RL-based motion-cadence steering method that encourages the model to explore and focus on sign-level keyframes. Experimental results on How2Sign and OpenASL benchmarks show that our proposed AttnSign consistently outperforms existing methods.

发表机构

  • Rochester Institute of Technology(罗切斯特理工学院)
  • University of Virginia(弗吉尼亚大学)
  • National Technical Institute for the Deaf(国家聋人技术学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑