arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SCOUT:面向超长时长自我中心视频推理的自检查与恢复感知工具思维智能体

SCOUT: Self-Checking and Recovery-Aware Tool-Thought Agents for Ultra-Long Egocentric Video Reasoning

Keyang Zhong, Kuo Wang, Peng Liu, Quanlong Zheng, Junlin Xie, Zhijia Liang, Yanhao Zhang, Guanbin Li

arXiv 2608.07959首次发表:更新:

发表机构

Sun Yat-sen University; Shenzhen Loop Area Institute; Guangdong OPPO Mobile Telecommunications Corp., Ltd.; OPPO AI Center; The Chinese University of Hong Kong, Shenzhen(中山大学; 深圳河套学院; 广东欧珀移动通信有限公司; OPPO人工智能中心; 香港中文大学(深圳))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出SCOUT智能体框架,结合自适应策略与UPS-GRPO方法,解决超长自我中心视频推理的错误传播及信用分配问题,在对应基准上达最优性能且短时长场景仍具竞争力。

AI 中文摘要

超长时长自我中心视频理解需要对分布在数小时或数天内的时间稀疏证据进行推理,这对当前上下文有限的多模态模型以及关键视频片段的定位提出了挑战。虽然工具思维链(Chain-of-Tool-Thought,CoTT)智能体系统支持迭代检索和检查,但由于缺乏恢复机制的刚性放大策略,它们会遭受错误传播。在本研究中,我们通过SCOUT(Self-Checking Chain-Of-Tool-thought,自检查工具思维链)解决这些挑战,这是一种恢复感知的智能体框架,引入了自适应策略,用于评估中间工具观测结果并动态权衡利用(放大)和探索(区域切换),从而能在极长的时间范围内实现稳健的多跳推理。然而,训练这种多轮使用工具的智能体仍然具有挑战性,因为现有的强化学习(RL)方法依赖于稀疏的结果级奖励,且缺乏对扩展决策轨迹的监督,导致长时序推理的信用分配次优。为解决这一问题,我们开发了UPS-GRPO,这是一种不确定性优先的策略优化方法,将探索集中在工具后高不确定性状态,同时保持样本效率。我们进一步引入了轮次级优势分解,将结果奖励与基于工具的时间对齐奖励相结合,以改善信用分配。实验表明,SCOUT在超长时长自我中心视频基准上达到了最先进的结果,同时在较短时间范围的长视频设置上仍具有竞争力。

英文摘要

Ultra-long egocentric video understanding requires reasoning over temporally sparse evidence distributed across hours or days, challenging current multimodal models with limited context and the grounding of key video segments. While Chain-of-Tool-Thought (CoTT) agent systems enable iterative retrieval and inspection, they suffer from error propagation due to rigid zoom-in strategies that lack recovery mechanisms. In this work, we address these challenges through SCOUT (Self-Checking Chain-Of-Tool-thought), a recovery-aware agentic framework introducing an adaptive policy that evaluates intermediate tool observations and dynamically trades off exploitation (zoom-in) and exploration (region switching), enabling robust multi-hop reasoning over extremely long horizons. However, training such multi-turn tool-using agents remains challenging, as existing RL methods rely on sparse outcome-level rewards and lack supervision over extended decision trajectories, resulting in suboptimal credit assignment for long-horizon reasoning. To address this, we develop UPS-GRPO, an uncertainty-prioritized policy optimization method that concentrates exploration on high-uncertainty post-tool states while preserving sample efficiency. We further introduce a turn-level advantage decomposition that integrates outcome rewards with tool-grounded temporal alignment rewards for improved credit assignment. Experiments show that SCOUT achieves state-of-the-art results on ultra-long egocentric benchmarks, while remaining competitive on shorter-horizon long-video settings.

CommentsAccepted by ACM Multimedia 2026 (MM '26)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑