手语问答:手语理解的新任务、基准与基线模型
Sign Language Question Answering: A New Task, Benchmark, and Baseline for Sign Language Understanding
浏览论文内容
中文总结 AI 辅助
该研究提出手语问答新任务,构建基于PHOENIX14T与CSL-Daily的SignQA基准,设计带特定模块的基线模型,实验显示其在各问题类别上优于代表性视觉-语言模型。
中文摘要 AI 辅助
近期手语理解(SLU)领域的进展已在连续手语识别、手语翻译等任务中取得显著成果。然而,这些任务均基于预设目标设计,要求模型学习从手语视频到词汇或口语句子的固定映射,因此仅能有限地评估模型是否真正理解手语视频的语义内容。为解决这一局限,我们首先提出手语问答(SLQA)这一新任务,该任务要求模型回答关于手语视频的任意自然语言问题,以此评估手语理解能力。与以往手语理解任务不同,SLQA提供更灵活全面的评估框架,可评估识别与翻译之外的多种推理能力。为支撑该任务,我们基于PHOENIX14T与CSL-Daily构建两个SignQA基准,通过精心设计的模板从现有词汇与句子标注自动生成问答对,生成的数据集涵盖位置推理、结构推理、视觉搜索、词汇识别、翻译理解五大互补问题类别。最后,我们提出一个简洁有效的基线模型,配备问题条件调制时间下采样模块与领域内知识迁移策略,可实现从现有手语理解任务的有效知识迁移,同时增强问题感知的时间特征建模。大量实验表明,该基线模型在所有问题类别上均持续优于代表性视觉-语言模型,为未来手语问答研究建立了坚实基准。数据集可在{ this https URL }获取。
英文摘要
Recent advances in sign language (SL) understanding (SLU) have led to remarkable progress in tasks such as continuous SL recognition and SL translation. However, these tasks are designed with predefined objectives, requiring models to learn a fixed mapping from sign videos to glosses or spoken-language sentences. As a result, they provide only a limited assessment of whether a model truly understands the semantic content of SL videos. To address this limitation, \textbf{we first propose a new task, Sign Language Question Answering (SLQA)}, which evaluates SL understanding by requiring models to answer arbitrary natural language questions about SL videos. Unlike previous SLU tasks, SLQA provides a more flexible and comprehensive evaluation framework that assesses multiple reasoning capabilities beyond recognition and translation. To facilitate this task, \textbf{we further construct two SignQA benchmarks} based on PHOENIX14T and CSL-Daily by automatically generating question-answer pairs from existing gloss and sentence annotations using carefully designed templates. The resulting datasets cover five complementary question categories, including position reasoning, structural reasoning, visual search, gloss recognition, and translation understanding. \textbf{Finally, we propose a simple yet effective baseline model} equipped with a Question-Conditioned Modulated Temporal Downsampling module and an in-domain knowledge transfer strategy, enabling effective knowledge transfer from existing SLU tasks while enhancing question-aware temporal feature modeling. Extensive experiments demonstrate that our baseline consistently outperforms representative vision-language models across all question categories, establishing a strong benchmark for future research on SLQA. Datasets are available at:{https://huggingface.co/datasets/hulala/SignQA-2026}.