arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.11899cs.CVcs.HC

一次标注,按需取帧:面向预算受限的智能体长视频理解的视觉需求路由

Caption-once, Frames-on-Demand: Visual-Need Routing for Budget-Aware Agentic Long Video Understanding

Weitong Cai, Hang Zhang, Yukai Huang, Yiqiao Xie, Shan Gao, Jiankang Deng, Songcen Xu, Jifei Song, Zhensong Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

针对边缘设备长视频理解,提出CFD框架,利用视觉-文本二元性,通过离线标注索引和按查询的视觉需求路由,仅对感知问题取帧,在降低计算成本的同时保持高精度。

中文摘要 AI 辅助

边缘设备上的长视频理解必须在紧凑的计算和带宽预算下对数小时的内容进行推理。对视觉令牌进行子采样会丢失时间结构,而纯文本的视频记忆则会丢失细粒度的视觉属性。我们观察到一种视觉-文本二元性:语言记忆比密集帧更能承载长程时间结构,而像素对于属性级感知仍然具有决定性作用。基于这一见解,我们提出了“一次标注,按需取帧”(CFD),一个预算感知的边缘-云智能体框架。边缘端执行一次离线标注过程,构建一个双轨叙事索引——事件级故事骨架加上片段级微日志——该索引被缓存并在查询间复用,无需重新标注。在查询时,云端多模态大语言模型(MLLM)以故事优先的循环对该索引进行推理,该循环以轻量级视觉需求路由器为中心:一个按查询的门控模块,仅对感知类问题(外观、屏幕文本、属性消歧)触发有界的关键帧检索,并将时间结构类问题保留在语言空间中。该路由器将视觉访问转化为一等公民的、由查询条件决定的成本,无论视频长度如何,都限制每次查询的帧消耗。在长视频基准上的实验表明,在显著减少在线视觉处理的同时,实现了强大的精度-效率权衡。

英文摘要

Long-video understanding on edge devices must reason over hours of content under tight compute and bandwidth budgets. Subsampling visual tokens loses temporal structure, while text-only video memories lose fine-grained visual attributes. We observe a visual-textual duality: language memories carry long-range temporal structure better than dense frames, while pixels remain decisive for attribute-level perception. Building on this insight, we propose Caption-once, Frames-onDemand (CFD), a budget-aware edge-cloud agentic framework. The edge runs a single offline captioning pass that builds a dual-track narrative index, an event-level story skeleton plus a clip-level micro-log, cached and reused across queries without re-captioning. At query time, a cloud-side MLLM reasons over the index in a story-first loop centered on a lightweight Visual-Need Router: a per-query gating module that triggers bounded keyframe retrieval only for perceptual questions (appearance, on-screen text, attribute disambiguation) and keeps temporal-structural questions in language space. The router turns visual access into a first-class, query-conditioned cost, capping per-query frame consumption regardless of video length. Experiments on long-video benchmarks demonstrate strong accuracy-efficiency trade-offs while substantially reducing online visual processing.

发表机构

  • Queen Mary University of London(伦敦玛丽女王大学)
  • Durham University(杜伦大学)
  • Imperial College London(帝国理工学院)
  • Huawei(华为)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑