发表机构
Central China Normal University; Wuhan United Imaging Healthcare Co., Ltd.(华中师范大学; 武汉联影医疗科技有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对教学视频检索中抽象概念与具体视觉证据的匹配难题,提出概念驱动域适应(CDDA)三阶段框架,通过概念锚点对齐文本与视觉空间,在中学物理基准上优于多模态基线。
AI 中文摘要
科学教师经常不是通过描述屏幕上出现的内容,而是通过查询他们打算教授的抽象概念来搜索纪录片片段。这一使用场景暴露了现有基于语言的视频时刻检索方法的局限性,这些方法通常假设查询描述的是可观察的事件,而教学搜索需要检索体现潜在科学原理的具体视觉现象。我们将这一设置研究为概念到示例的视频检索,这是一个在干草堆中寻找抽象针的问题,其中紧凑的课程概念必须基于时间上稀疏的纪录片证据。为了弥合这一抽象差距,我们提出了概念驱动域适应(CDDA),这是一个三阶段框架,用于将双塔视觉-语言模型适应到概念级检索。CDDA将概念视为中间语义锚点:它首先用教科书和教师手册的示例-概念对构建文本嵌入空间,然后在冻结的视觉编码器下将这种概念感知的几何结构转移到纪录片视觉内容,最后通过稀疏的视觉概念监督联合适应两个编码器。从几何角度来看,这种分阶段对齐减少了文本-概念和视觉-概念的角度差距,从而鼓励概念级适应,同时保留预训练模型的具体图像描述对齐。在一个精心策划的中学物理检索基准上,CDDA在面向教学的概念检索方面优于包括Qwen3-VL-Embedding-2B在内的几个竞争性多模态基线,同时在适应后保持了具体的图像-文本匹配。
英文摘要
Science teachers frequently search for documentary excerpts not by describing what appears on screen, but by querying the abstract concepts they intend to teach. This use case exposes a limitation of existing language-based video moment retrieval methods, which typically assume that queries describe observable events, whereas instructional search requires retrieving concrete visual phenomena that instantiate an underlying scientific principle. We study this setting as concept-to-example video retrieval, an abstract-needle-in-a-haystack problem where compact curriculum concepts must be grounded in temporally sparse documentary evidence. To bridge this abstraction gap, we propose Concept-Driven Domain Adaptation (CDDA), a three-stage framework for adapting two-tower vision-language models to concept-level retrieval. CDDA treats concepts as intermediate semantic anchors: it first structures the textual embedding space with textbook and teacher-handbook example-concept pairs, then transfers this concept-aware geometry to documentary visuals under a frozen visual encoder, and finally jointly adapts both encoders with sparse visual concept supervision. From a geometric perspective, this staged alignment reduces text-concept and vision-concept angular gaps, thereby encouraging concept-level adaptation while preserving the pretrained model's concrete image description alignment. On a curated middle-school physics retrieval benchmark, CDDA achieves stronger pedagogically oriented concept retrieval than several competitive multimodal baselines, including Qwen3-VL-Embedding-2B, while maintaining concrete image-text matching after adaptation.
Comments19 pages, 10 figures