arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.01980cs.CVcs.AI

AdaThinkV:面向令牌高效视频推理的自适应思维

AdaThinkV: Adaptive Thinking for Token-Efficient Video Reasoning

Jingqi Tian, Haoji Zhang, Lin Chen, Hongbo Jin, Haonan Xu, Tianrui Zhu, Xingming Shui, Shilin Ma, Wenjing Yang, Yansong Tang

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出AdaThinkV自适应视频推理框架,通过ThinkGain与VRPO解决CoT推理的令牌浪费问题,在视频推理任务中以更少令牌实现优于最强自适应基线的准确率。

中文摘要 AI 辅助

思维链(CoT)推理可提升复杂视频问题的性能,但常对简单问题浪费解码令牌。本研究探讨视频多模态大语言模型是否可针对每个问题调整推理力度,提出AdaThinkV,一种无需离线难度标签、人工调优置信阈值或外部路由器的自适应视频推理框架。强化学习阶段,AdaThinkV为每个提示在显式推理与直接回答模式中采样匹配的rollout;ThinkGain通过权衡显式推理的准确率增益与额外响应长度,估计提示级效用,为条件响应生成与自主模式选择提供监督。对于困难提示,有限的rollout探索会产生每组响应均不成功、准确率奖励变化极小的情况,学习信号不足,因此引入方差恢复策略优化(VRPO),保留并逐步扩展这些组,从困难但可解的提示中恢复有效信号。推理阶段,AdaThinkV选择响应模式并以单一自回归序列生成响应。在统一视频推理评估套件中,AdaThinkV实现平均准确率40.79,平均输出令牌257.20,较评估的最强自适应基线准确率高2.98个百分点,令牌使用量减少22.7%。项目页面:this https URL

英文摘要

Chain-of-thought (CoT) reasoning can improve performance on difficult video questions but often wastes decoding tokens on simple ones. We study whether a video multimodal large language model can adapt its reasoning effort to each question. We propose AdaThinkV, an adaptive framework for video reasoning that learns whether to reason explicitly without offline difficulty labels, manually tuned confidence thresholds, or an external router. During reinforcement learning, AdaThinkV samples matched rollouts in explicit reasoning and direct answering modes for each prompt. ThinkGain estimates the prompt-level utility of explicit reasoning by balancing its accuracy gain against additional response length, providing supervision for both conditional response generation and autonomous mode selection. For difficult prompts, limited rollout exploration can yield groups in which every response is unsuccessful and accuracy rewards show little variation, providing insufficient signal for learning. We therefore introduce Variance Recovery Policy Optimization (VRPO), which retains and progressively expands these groups to recover informative signals from prompts that are difficult yet solvable. At inference, AdaThinkV selects a response mode and generates the response in a single autoregressive sequence. Across a unified suite of video reasoning evaluations, AdaThinkV achieves a mean accuracy of 40.79 with an average of 257.20 output tokens, outperforming the strongest evaluated adaptive baseline by 2.98 points while using 22.7% fewer tokens. Project page: https://trilarflagz.github.io/AdaThinkV/

↑