arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CausalChapter:通过干预依赖建模改进长视频章节划分

CausalChapter: Improving Long-Video Chaptering with Interventional Dependency Modeling

Xinran Duan, Guozhang Li, Yaoyao Zhong, Mei Wang, Lizhi Wang, Hua Huang

arXiv 2609.08686首次发表:更新:

发表机构

School of Artificial Intelligence, Beijing Normal University; Beijing Key Laboratory of Artificial Intelligence for Education; Engineering Research Center of Intelligent Technology and Educational Application, Ministry of Education(北京师范大学人工智能学院; 北京市教育人工智能重点实验室; 教育部智能技术与教育应用工程研究中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

CausalChapter提出一种干预式依赖建模框架,通过局部依赖偏移和跨段支持选择,解决长视频章节划分中的边界错误传播和上下文碎片化问题,提升边界定位与描述质量。

AI 中文摘要

长格式教学视频需要自动章节划分以支持浏览、导航和知识访问。最近的长上下文语言模型可以从文本化的视频输入执行章节划分,但对于内容密集、转录文本较长、主题过渡平滑且章节输出详细的讲座视频,它们仍然成本高昂且脆弱。一种可扩展的“分段-然后-描述”范式降低了这一成本,但引入了两个新挑战:边界错误传播和跨章节上下文碎片化。我们提出了CausalChapter,一个受干预启发的长视频章节划分框架,通过轻量级掩蔽和移除干预来估计预测级别的影响。对于边界定位,我们的局部依赖偏移模块检测相邻时间窗口之间预测依赖性的下降;对于章节描述生成,我们的跨段支持选择模块根据历史上下文对当前预测的支持程度对其重新排序。在长视频章节划分基准上的实验表明,CausalChapter提高了边界定位、章节描述质量和跨章节连贯性。

英文摘要

Long-form instructional videos require automatic chaptering to support browsing, navigation, and knowledge access. Recent long-context language models can perform chaptering from textualized video inputs, but they remain costly and brittle for content-dense lecture videos with long transcripts, smooth topic transitions, and detailed chapter outputs. A scalable segment-then-caption paradigm reduces this cost, but introduces two new challenges: boundary error propagation and fragmented cross-chapter context. We propose \textbf{CausalChapter}, an intervention-inspired framework for long-video chaptering that estimates prediction-level influence through lightweight masking and removal interventions. For boundary localization, our Local Dependency Shift module detects drops in predictive dependency between adjacent temporal windows; for chapter description generation, our Cross-Segment Support Selection module reranks historical contexts according to their support for the current prediction. Experiments on long-video chaptering benchmarks show that CausalChapter improves boundary localization, chapter description quality, and cross-chapter coherence.

CommentsAccepted by EMNLP 2026 conference

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑