AI 中文总结
本文提出新的音频描述生成范式,构建含206部电影的REFRAMED数据集及评估协议,实验显示现有AD系统和多模态LLMs表现优于简单基线但不及专业人类,为视频理解研究提供新基础。
AI 中文摘要
音频描述(Audio Description,AD)是对视频中关键视觉内容的口头叙述,旨在让视障观众获取相关信息。与标准视频字幕不同,AD是一项结构化的编辑任务:描述必须插入到对话间隙中,且仅需传达理解叙事所需的内容。然而,现有方法将AD生成置于人工设定场景中,描述的内容和时间均预先指定,使任务简化为片段级字幕生成;此外,这些方法还依赖有噪声的转录与对齐流程,且缺乏建模叙事上下文所需的丰富平行数据。本文提出一种新的AD生成范式,要求模型需同时决定描述什么以及何时描述。为支撑该范式,我们推出REFRAMED数据集,包含206部电影的3302个场景,共2023个视频,配备专业AD转录本(含美式与英式版本)、专业字幕及对齐剧本;还提供人工整理的挑战集,将完整电影与多个AD参考配对,以及利用对话间隙和多参考比较的评估协议。对最先进的AD系统和多模态大语言模型(multimodal LLMs)的实验表明,它们虽优于简单基线,但远未达到专业人类表现。我们的数据集与基准为视频理解研究奠定了新基础。
英文摘要
Audio Description (AD) is a verbal narration of key visual content in videos, enabling access for visually impaired audiences. Unlike standard video captioning, AD is a structured editorial task: descriptions must be inserted into gaps in dialogue and must convey only what is needed to understand the narrative being told. However, existing approaches formulate AD generation in an artificial setting where both the content and timing of descriptions are pre-specified, reducing the task to clip-level captioning. They further rely on noisy transcription and alignment pipelines, and lack the rich parallel data required for modeling narrative context. We introduce a new formulation of AD generation in which models must jointly decide what to describe and when to do it. To support this, we present REFRAMED, a high-quality dataset of 2,023 videos that span 3,302 scenes from 206 movies, with professional AD transcripts (both American and British versions), professional subtitles and aligned screenplays. We also provide a manually curated challenge set that pairs full movies with multiple AD references, together with evaluation protocols that leverage dialogue gaps and multi-reference comparisons. Experiments with state-of-the-art AD systems and multimodal LLMs show that they outperform trivial baselines but fall far short of expert human performance. Our dataset and benchmark establish a new foundation for research on video understanding.
CommentsCOLM 2026