发表机构
School of Computer Science and Technology, Beijing Institute of Technology; Department of Computing, The Hong Kong Polytechnic University; Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences(北京理工大学计算机科学与工程学院; 香港理工大学计算学系; 中国科学院深圳先进技术研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
NLPCC 2026的DA - MIVQA共享任务,通过区分问题难度扩展多语言多模态医学视频基准,含三个赛道。其数据集来自公共医学教学渠道,涵盖多种场景并经人工标注难度。该任务为评估医学教学视频问答系统提供实用基准。
AI 中文摘要
继NLPCC 2023 - 2025中的CMIVQA、MMI - VQA和M4IVQA挑战之后,我们为NLPCC 2026引入了难度感知医学教学视频问答(DA - MIVQA)共享任务。DA - MIVQA通过根据回答所需证据的类型和复杂性明确区分问题,扩展了先前的多语言多模态医学视频基准。具体而言,简单问题通常可从基于字幕的文本线索回答,复杂问题则需要视觉定位、程序理解和跨模态证据整合。该挑战包含三个赛道。数据集从公共医学教学渠道收集,涵盖多种场景并带有难度注释。本文介绍了DA - MIVQA的任务动机、数据集构建、评估协议、参与概况、竞赛结果和代表性系统。DA - MIVQA为评估不同文本、视觉、时间和程序推理要求下的医学教学视频问答系统提供了实用基准。
英文摘要
Following the CMIVQA, MMI-VQA, and M4IVQA challenges in NLPCC 2023--2025, we introduce the Difficulty-Aware Medical Instructional Video Question Answering (DA-MIVQA) shared task for NLPCC 2026. DA-MIVQA extends previous multilingual and multimodal medical video benchmarks by explicitly distinguishing questions according to the type and complexity of evidence required for answering. Specifically, simple questions can often be answered from subtitle-based textual cues, whereas complex questions require visual grounding, procedural understanding, and cross-modal evidence integration. The challenge contains three tracks: Difficulty-Aware Temporal Answer Grounding in Single Video (DA-TAGSV), Difficulty-Aware Video Corpus Retrieval (DA-VCR), and Difficulty-Aware Temporal Answer Grounding in Video Corpus (DA-TAGVC). The dataset is collected from public medical instructional channels, covers diverse scenarios such as first aid, emergency response, rehabilitation, nursing, and general medical education, and is manually verified with difficulty annotations. This paper presents the task motivation, dataset construction, evaluation protocol, participation overview, competition results, and representative systems of DA-MIVQA. DA-MIVQA provides a practical benchmark for evaluating medical instructional video question answering systems under varying textual, visual, temporal, and procedural reasoning requirements.
Comments19 pages, 6 figures, 6 tables