arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向未修剪自我中心视频的预算受限视觉-语言字幕生成的预解码音频分诊

Pre-Decoding Acoustic Triage for Budgeted Vision-Language Captioning of Untrimmed Egocentric Video

Masoud Jalayer, Changyi Li, Yu Xiao

arXiv 2608.22359首次发表:更新:

发表机构

Aalto University(阿尔托大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对预算受限的未修剪自我中心视频视觉-语言字幕生成问题,提出预解码音频分诊策略,通过音频优先选择窗口减少VLM调用,提升动作覆盖率,在多个数据集上优于现有方法。

AI 中文摘要

自动分析时长数小时的自我中心视频,对于物流、建筑和制造业的进度监控、质量控制及安全而言愈发重要。然而,当前采用视觉-语言模型(VLM)处理短的固定尺寸窗口的流程成本极高,因为成本随模型调用次数增加而上升。为降低该成本,现有研究提出了分诊策略来选择值得调用VLM的窗口,但这些策略要么均匀采样,要么使用视觉特征对窗口排序,讽刺的是这需要进行视频解码,而这正是预算约束旨在避免的操作。我们提出音频优先分诊:采用最轻量的模态(音频)选择窗口,在解码任何视频帧前进行评分,因此该方法可自然地与 token 压缩或量化相结合。其创新点在于目标而非表示:我们训练选择器并非用于逐帧声音事件检测,而是每触发一次对应一个动作。在所有评估的调用率下,该目标调整使动作覆盖率提升了4.0至10.8个百分点,所用特征为冻结的AudioSet预训练特征,无需领域特定的声音事件标签。在EPIC-KITCHENS-100(EK-100)数据集上,使用不足一半的可用调用时,该分诊策略在匹配覆盖率的情况下减少了9%至20%的VLM调用;在Ego4D数据集的247个视频片段上,其中段性能优于均匀采样;且优于两个近期的视觉关键帧选择器。代码、参考实现及本手稿引用的所有结果文件均位于此httpsURL。

英文摘要

Automatically analyzing hours-long egocentric video is increasingly essential for progress monitoring, quality control, and safety in logistics, construction, and manufacturing. Yet current pipelines that process short, fixed-size windows with a vision-language model (VLM) are prohibitively expensive because cost scales with the number of model calls. To reduce this cost, prior work proposes triage policies to select which windows merit a VLM invocation. However, these policies either sample uniformly or rank windows using visual features, which ironically requires the video decoding that the budget constraints are meant to avoid. We propose audio-first triage: select windows using the lightest modality, scored before any video frame is decoded, so the approach composes naturally with token compression or quantization. The novelty lies in the objective, not the representation: rather than a per-frame sound-event detector, we train the selector to trigger once per action. This objective shift improves action coverage by 4.0-10.8 percentage points across all evaluated call rates, using frozen AudioSet-pretrained features without domain-specific sound-event labels. Using fewer than half of the available calls, the triage cuts 9-20% of VLM calls at matched coverage on EPIC-KITCHENS-100 (EK-100), surpasses uniform sampling through the mid-range on Ego4D over 247 clips, and outperforms two recent visual keyframe selectors. Code, the reference implementation and every results file this manuscript reads are at https://github.com/masjalayer/PreDecoding-AcousticTriage.

Comments18 pages, 5 figures, 9 tables. Under review at IEEE BigData 2026, Industrial and Government Track. Code: https://github.com/masjalayer/PreDecoding-AcousticTriage

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑