arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LEAP:面向长音视频感知的分块证据检索

LEAP: Learned Block-wise Evidence Retrieval for Long Audio-Video Perception

Juyi Lin, Zhiqiang Lao, Jiali Cui, Lin Zhao, Pu Zhao, Dichang Zhang, Arman Akbari, Yu Qi, Xinru Jiang, Yanzhi Wang, Heather Yu, Liang Peng

arXiv 2609.39938首次发表:更新:

发表机构

Northeastern University; Futurewei Technologies(东北大学; 未来网络技术公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

LEAP通过分块检索和双LoRA训练,在长音视频问答中实现与时长无关的上下文,显著提升多个基准的性能。

AI 中文摘要

小时级别的音视频问答受制于上下文困境:对整段录音进行密集编码会迅速耗尽上下文限制,而均匀的时间压缩则会严重稀释细粒度的声学和视觉证据。我们提出了LEAP框架,该框架让模型自行检索证据,而无需将整段录音置于一个上下文中。LEAP将录音划分为固定时长的块,对每个块应用轻量级定位过程,以对短候选窗口进行评分。得分最高的窗口被汇集并在一次有界的答案传递中重新编码。因此,答案输入和峰值上下文与录音时长无关。通过将证据定位与推理解耦,我们的框架可以在预计算转录本上定位候选时间窗口,而无需解码媒体帧,同时通过将最终答案传递路由到原始音视频流上,保留细粒度的视觉和非语音证据。LEAP训练两个阶段:定位LoRA改进所选窗口,答案LoRA改进从相同窗口读取的答案。块网格天然支持因果查询,使LEAP能够支持流式推理而无需特定于流式的训练。在多个AVQA基准上,LEAP相对于Qwen3-Omni-30B-A3B基线提升了4.5%至16.8%,并迁移到第二个全模态骨干MiniCPM-o 4.5,超越其已发表结果3.1%至13.0%。

英文摘要

Hour-scale audio-visual question answering is constrained by a context dilemma: dense whole-recording encoding rapidly exhausts context limits, whereas uniform temporal compression severely dilutes fine-grained acoustic and visual evidence. We introduce LEAP, a framework where the model retrieves its own evidence without placing the whole recording in one context. LEAP divides a recording into fixed-duration blocks, applying a lightweight localization pass to each block to score short candidate windows. The highest-ranked windows are pooled and re-encoded in a single bounded answer pass. Consequently, the answer input and peak context remain independent of the recording duration. By decoupling evidence localization from reasoning, our framework can localize candidate temporal windows over pre-computed transcripts without decoding media frames, while preserving fine-grained visual and non-speech evidence by routing the final answering pass over raw audio-visual streams. LEAP trains both stages: a localization LoRA improves the selected windows, and an answer LoRA improves the answers read from the same windows. The block grid natively supports causal queries, enabling LEAP to support streaming inference without streaming-specific training. Across several AVQA benchmarks, LEAP improves over the Qwen3-Omni-30B-A3B baseline by 4.5-16.8%, and transfers to a second omni-modal backbone, MiniCPM-o 4.5, surpassing its published results by 3.1-13.0%.

Comments39 pages, 16 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑