arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.12549cs.CV

StrAD:面向长视频音频描述生成的流式方法与基准

StrAD: A Streaming Method and Benchmark for Audio Description Generation for Long-form Videos

  • Fraunhofer IAIS(弗劳恩霍夫智能分析与信息系统研究所)
  • University of Bonn(波恩大学)
  • University of Applied Sciences Bonn-Rhein-Sieg(波恩-莱茵-锡格应用科学大学)
  • Lamarr Institute for Machine Learning and Artificial Intelligence(拉马尔机器学习与人工智能研究所)
  • Center for Robotics, University of Bonn(波恩大学机器人中心)

机构由 AI 辅助整理,请以论文原文为准。

Julian Spravil, Sebastian Houben, Sven Behnke

AI总结:

针对长视频音频描述生成的需求,研究提出首个流式方法StrAD,构建了涵盖多类型全长视频的基准,其微调模型在相关任务上达到先进性能,为扩大无障碍覆盖提供支持。

AI中文摘要:

视觉内容是主要的传播媒介,但缺少音频描述(ADs)时,视障和低视力人群无法访问这些内容。音频描述会在自然的音频停顿间隙,叙述与上下文相关的视觉事件。手动制作音频描述成本高昂,导致其仅能覆盖少量可用内容。多数现有自动音频描述生成方法将该任务建模为视频片段字幕生成,需要真实时间戳以及角色数据库等额外上下文线索。当前基准也强化了这一建模方式,由短视频片段搭配自动或任务不匹配的标注构成。我们推出StrAD,这是针对涵盖电影、纪录片、短片、表演、电子游戏等多种类型的全长视频的长音频描述生成基准。我们将音频描述生成重新建模为流式密集视频字幕生成。我们的方法使用滑动窗口处理全长视频,在不依赖真实时间戳的情况下将音频描述插入现有字幕,支持微调模型以及视觉语言模型的零样本提示。在给定时间戳的片段级任务中,我们微调后的StrAD-FT在CMD-AD上达到36.3的CIDEr值(比Shot-by-shot高10.0),在StrAD上建立了51.0的CIDEr参考点,在MAD-Eval上以24.9的CIDEr值保持竞争力。在全长视频流式任务中,StrAD-FT的SODA得分为2.4,而我们的零样本基线StrAD-Zero为1.1,尽管两者在时间定位和叙事连贯性方面仍存在局限。此前研究以离线多阶段方式处理全长音频描述生成,而我们的方法是首个流式方法,无需真实时间戳即可实时生成音频描述。StrAD推进了全长音频描述生成的可衡量性,这是扩大无障碍覆盖范围的前提。

英文摘要:

Visual content is the dominant medium of communication, yet without audio descriptions (ADs), it remains inaccessible to blind and low-vision people. ADs narrate context-relevant visual events during natural audio pauses. Manually creating ADs is expensive, limiting coverage to a small fraction of available content. Most existing automatic AD generation methods frame the task as video clip captioning, requiring ground-truth timestamps and additional context cues such as character databases. Current benchmarks reinforce this framing, consisting of short video segments paired with automatic or task-mismatched annotations. We introduce StrAD, a benchmark for long-form AD generation on full-length videos spanning diverse genres such as movies, documentaries, short films, performances, and video games. We reformulate AD generation as streaming dense video captioning. Our approach processes full-length videos with a sliding window, inserting ADs into existing transcripts without ground-truth timestamps, and supports both fine-tuned models and zero-shot prompting of vision-language models. On the segment-level task with given timestamps, our fine-tuned StrAD-FT sets the state of the art on CMD-AD with 36.3 CIDEr (+10.0 over Shot-by-shot), establishes a reference point on StrAD (51.0 CIDEr), and remains competitive on MAD-Eval at 24.9 CIDEr. On the full-video streaming task, StrAD-FT reaches a SODA score of 2.4 against 1.1 for our zero-shot baseline StrAD-Zero, though both exhibit limitations in temporal localization and narrative coherence. While prior work has tackled full-video AD generation in an offline, multi-stage fashion, ours is the first streaming approach, generating ADs on the fly without ground-truth timestamps. StrAD makes progress on full-video AD generation measurable, a prerequisite for scaling accessibility.

↑