发表机构
Dartmouth College(达特茅斯学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对长格式音频描述问题,提出免训练的StoryTeller框架。它通过维护叙事记忆跨场景传递信息,仅用原始视频和标题,经语义过滤等确保信息准确。引入StoryAD-QA基准测试,实验证明该框架显著提升了叙事相关能力。
AI 中文摘要
长格式音频描述(AD)不仅要描述可见动作,还需保留跨场景的角色、事件、关系和故事情节,以便盲人和低视力(BLV)观众能理解电影。现代视频语言模型(VLMs)在短视频片段上有效,但处理长格式时往往独立对待每个时刻,导致描述缺失关键信息。本文提出StoryTeller,一个免训练的长格式AD框架。它通过维护经过验证的叙事记忆来跨场景传递与故事相关的信息,使后续描述保持连贯、有根据且上下文丰富。该方法无需字幕、脚本等,仅利用原始视频和电影标题,可检索公共电影元数据并通过语义过滤和VLM验证确保信息准确。为评估生成的AD是否保留叙事信息,引入了StoryAD-QA基准测试。实验表明,StoryTeller在自动、基于QA和人工评估中均显著提高了叙事连贯性、事实依据和故事理解能力。
英文摘要
Long-form audio description (AD) requires more than describing visible actions: it must preserve characters, events, relationships, and story context across scenes so that blind and low-vision (BLV) audiences can follow a film. Modern video-language models (VLMs) are effective on short clips, but they often treat each moment independently, producing descriptions that miss who characters are, why events matter, and how the current scene connects to earlier narrative context. We propose StoryTeller, a training-free framework for story-aware long-form AD. Instead of relying only on local visual cues, StoryTeller maintains a verified narrative memory that carries forward story-relevant information across scenes, enabling later descriptions to remain coherent, grounded, and contextually informative. Given only raw video and a movie title, StoryTeller can optionally retrieve public movie metadata to resolve names and story context, while accepting only facts that are supported by the video through semantic filtering and VLM verification. The method requires no subtitles, scripts, AD transcripts, aligned captions, character banks, precomputed face identities, or task-specific fine-tuning. To evaluate whether generated AD preserves narrative information, we introduce StoryAD-QA, a question-answering benchmark that tests whether a language model can answer story-context questions using only the generated descriptions. Experiments on standard AD benchmarks and diverse long-form videos show that StoryTeller consistently improves narrative coherence, factual grounding, and story comprehension over strong baselines in automatic, QA-based, and human evaluations.
CommentsAccepted to the European Conference on Computer Vision (ECCV) 2026