什么、何时以及如何:将音频描述视为受约束的全局优化
What, When, and How: Audio Description as Constrained Global Optimization
AI总结:
本文提出将音频描述生成视为受约束的全局优化问题,利用LLM和混合整数线性规划联合决策描述内容、时间与表述,在REFRAMED基准上达到新SOTA。
AI中文摘要:
音频描述(AD)通过在对话间隙叙述视觉信息,使电影对盲人和视力受损观众可及。现有的自动AD系统大多将生成视为局部视频到文本问题,假设要描述的内容及其时间位置已经给定。而现实中的AD则需要对哪些视觉信息在叙事上重要、何时可以叙述而不干扰对话、以及如何表述以适应可用时间这三个方面进行耦合决策。我们将AD生成形式化为一个关于这三个决策的受约束优化问题。我们的混合系统使用大型语言模型来提出并锚定视觉元素,估计它们对叙事的显著性,并生成压缩的实现。随后,一个混合整数线性规划在场景中联合选择并调度描述,受时间约束限制。在REFRAMED(一个用于现实电影AD的基准)上评估时,我们的方法在描述什么和何时描述方面比提示的LLM做出更好的决策,在叙事问答和时间接地指标上建立了新的最先进水平。消融研究表明,显式时间约束驱动了放置方面的改进,而显著性估计控制着保留多少叙事有用的内容。改进集中在时间和叙事指标上,而非n-gram重叠,尽管与专业描述者之间仍存在显著差距。
英文摘要:
Audio Description (AD) makes movies accessible to blind and visually impaired audiences by narrating visual information in gaps between dialogue. Existing automatic AD systems largely treat generation as a local video-to-text problem, assuming that the content to describe and its temporal location are already provided. Realistic AD instead requires coupled decisions about what visual information is narratively important, when it can be spoken without interfering with dialogue, and how it should be formulated to fit within the available time. We formalize AD generation as a constrained optimization problem over these three decisions. Our hybrid system uses large language models to propose and ground visual elements, estimate their salience to the narrative, and generate compressed realizations. A mixed-integer linear program then jointly selects and schedules descriptions across a scene subject to temporal constraints. When evaluated on REFRAMED, a benchmark for realistic AD of movies, our approach makes better decisions than prompted LLMs about what to describe and when to describe it, establishing a new SOTA on narrative QA and temporally grounded metrics. Ablations show that explicit temporal constraints drive gains in placement, while salience estimation controls how much narratively useful content is retained. Improvements are concentrated on temporal and narrative measures rather than n-gram overlap, although a significant gap to professional describers remains.