arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.03742cs.CVcs.CL

KnowVis:面向视频讲座的以知识为中心的视觉摘要

KnowVis: Knowledge-Centric Visual Summarization for Video Lectures

  • City University of Hong Kong(香港城市大学)
  • ETH Zürich(苏黎世联邦理工学院)

机构由 AI 辅助整理,请以论文原文为准。

Yi Xu, Yifan Hou, Xiaoyu Zhang

中文总结 AI 辅助

针对视频讲座线性信息传递与人类学习需构建认知网络的不匹配问题,提出KnowVis框架,生成精准清晰的视觉摘要,降低认知负荷并提升学习效果,还发布了配套教育视频数据集。

中文摘要 AI 辅助

视频讲座是极具价值的教育资源,但其密集且冗长的格式常令初学者难以承受。这一困难源于根本的教学不匹配:视频以线性方式传递瞬时信息,而人类学习需要构建相互关联的认知网络,对于缺乏先验领域知识的初学者而言,该任务会引发严重的认知负荷。现有视频摘要方法无法解决这种不匹配,因为它们主要生成以文本为主的线性摘要,仍需较高的认知努力。为弥合这一差距,我们提出KnowVis,一个将线性视频讲座转化为基于教学原理的视觉叙事的框架。KnowVis首先从多模态视频内容中提取详细的概念图,以识别重要且具有挑战性的阈值概念,随后构建结构化知识单元,最终合成引人入胜的视觉摘要。除该框架外,我们还引入了包含10个学科的125个教育视频的精选数据集,搭配1079个生成的视觉摘要。大量自动评估和人类研究表明,与最先进的基线方法相比,KnowVis生成更准确清晰的视觉内容,成功降低认知负荷,并显著提升学生的学习效果和知识保留率。

英文摘要

Video lectures are valuable educational resources, but their dense and lengthy formats often overwhelm novice learners. This difficulty stems from a fundamental pedagogical mismatch: while videos deliver transient information linearly, human learning requires constructing interconnected cognitive networks, a task that induces severe cognitive overload for novice learners lacking prior domain knowledge. Existing video summarization methods fail to resolve this mismatch, as they primarily produce text-heavy, linear condensations that still demand high cognitive effort. To bridge this gap, we propose KnowVis, a framework that transforms linear video lectures into pedagogically grounded visual narratives. KnowVis first extracts a detailed concept map from multimodal video content to identify important and challenging threshold concepts, then constructs structured knowledge units, and finally synthesizes engaging visual summaries. Alongside the framework, we introduce a curated dataset of 125 educational videos across 10 academic disciplines, paired with 1,079 generated visual summaries. Extensive automated evaluations and a human study demonstrate that, compared to state-of-the-art baselines, KnowVis generates more accurate and clear visuals that successfully reduce cognitive load and significantly improve student learning effectiveness and knowledge retention.

补充信息

相关深度报道

↑