arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LAVE:面向视频工具使用智能体的隐式视觉证据增强规划

LAVE: Latent Visual Evidence-Enhanced Planning for Video Tool-use Agents

Zijian Wang, Junnan Zhu, Rongzhen Li, Xiao Liu, Guohui Xiang, Quan Lu, Lijia Liu, Yining Wang, Jiang Zhong, Kaiwen Wei

arXiv 2608.07585首次发表:更新:

AI 中文总结

针对视频工具使用智能体的工具观测瓶颈,提出无训练框架LAVE,通过双通道观测接口复用隐式视觉证据,在三个基准数据集上提升了智能体性能,Video-MME得分较最强基线提高3.76分。

AI 中文摘要

长视频理解要求模型能从冗长且冗余的视频流中高效获取并复用稀疏的视觉证据。近期的视频工具使用智能体通过在不同时间尺度上迭代调用视觉工具来应对这一挑战,但它们的工具-规划器通信通常依赖文本观测。这种仅文本的接口会提供工具计算的有损摘要,导致之前计算的未被文本化的视觉证据被丢弃,无法用于后续规划。我们将这一局限称为工具观测瓶颈,并提出了无训练框架LAVE(Latent Visual Evidence-Enhanced Planning,隐式视觉证据增强规划),用于复用已完成工具调用的隐式视觉证据。LAVE引入了双通道观测接口:可见通道保留原始文本轨迹,隐式通道存储预文本化的视觉更新及其工具角色、源帧时间戳和视觉位置。规划期间,LAVE会检索与当前规划器状态相关但未被文本观测覆盖的证据,并通过带熵约束的帧时间路由、有界时间戳对齐的隐式更新来整合这些证据。这使视频智能体无需额外训练、帧重放或对原始编排进行修改即可复用现有视觉计算。在Video-MME、LongVideoBench和CG-Bench上的大量实验表明,LAVE在不同主干网络下均能持续提升视频工具使用智能体的性能。在可比的帧预算下,LAVE比最强基线将Video-MME总体得分提高了3.76分,证明了隐式视觉证据复用于多步骤视频智能体规划的有效性。

英文摘要

Long-video understanding requires models to efficiently acquire and reuse sparse visual evidence from long and redundant video streams. Recent video tool-use agents address this challenge by iteratively invoking visual Tools at different temporal scales, but their Tool-Planner communication typically relies on textual observations. Such text-only interfaces provide lossy summaries of Tool computations, causing previously computed visual evidence not verbalized to be discarded and unavailable for subsequent planning. We identify this limitation as the Tool observation bottleneck and propose Latent Visual Evidence-Enhanced Planning (LAVE), a training-free framework for reusing latent visual evidence from completed Tool calls. LAVE introduces a dual-channel observation interface: the visible channel preserves the original textual trajectory, while the latent channel stores pre-verbal visual updates with their Tool roles, source-frame timestamps, and visual locations. During planning, LAVE retrieves evidence relevant to the current Planner state but not covered by textual observations, and integrates it through bounded timestamp-aligned latent updates with entropy-constrained frame-time routing. This enables video agents to reuse existing visual computation without additional training, frame replay, or modifications to the original orchestration. Extensive experiments on Video-MME, LongVideoBench, and CG-Bench show that LAVE consistently improves video tool-use agents across backbones. Under a comparable frame budget, LAVE improves the Video-MME overall score by 3.76 points over the strongest baseline, demonstrating the effectiveness of latent visual evidence reuse for multi-step video-agent planning.

Comments16 pages, 6 figures, 9 tables. Includes appendix

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑