arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向视觉语言推理的令牌解耦潜在测试时缩放

Token-Disentangled Latent Test-Time Scaling for Vision-Language Reasoning

Hao-Xuan Ma, Yihao Liu, Yutao Sun, Yanting Miao, Mengyu Zhou, YiCheng Xiao, Long Chen, Zhenguo Li, Han-Jia Ye, Xiaoxi Jiang, Guanjun Jiang

arXiv 2609.35228首次发表:更新:

发表机构

Qwen Business Unit of Alibaba; School of Artificial Intelligence, Nanjing University; National Key Laboratory for Novel Software Technology, Nanjing University; Zhejiang University; University of Waterloo; Chinese Academy of Sciences; The Hong Kong University of Science and Technology; Frontier Robotics(阿里巴巴通义千问事业部; 南京大学人工智能学院; 南京大学计算机软件新技术全国重点实验室; 浙江大学; 滑铁卢大学; 中国科学院; 香港科技大学; 前沿机器人)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对多模态大语言模型推理中潜在测试时缩放忽略令牌角色差异的问题,提出令牌解耦的潜在测试时缩放框架,通过路由视觉和推理反馈至不同令牌,在Qwen2.5-VL-7B和InternVL3.5-8B上分别提升宏观准确率+2.57和+1.51。

AI 中文摘要

潜在测试时缩放通过推理过程中优化隐藏状态来提升推理能力,但现有方法通常对所有可编辑的潜在令牌应用单一的标量奖励。对于多模态大语言模型,这种全局更新忽略了生成的令牌扮演不同角色:有些对视觉证据敏感,而另一些则对应不确定的推理决策。我们提出了Token-Disentangled Latent Test-Time Scaling,一个推理时框架,使潜在优化具有令牌角色感知能力。从初始生成的轨迹开始,我们优化一个短的隐藏状态前缀,同时将感知侧的视觉反馈路由到图像敏感令牌,将推理反馈路由到高熵令牌。未被任一路由选中的令牌由锚点正则化器约束。在Qwen2.5-VL-7B和InternVL3.5-8B上的感知和推理基准测试中,我们的方法将宏观准确率分别比CoT提高了+2.57和+1.51,并且在匹配解码候选预算下优于强输出空间测试时缩放基线。代码可在该https URL获取。

英文摘要

Latent test-time scaling improves reasoning by refining hidden states during inference, but existing methods typically apply a single scalar reward to all editable latent tokens. For multimodal large language models, this global update ignores that generated tokens play different roles: some are sensitive to visual evidence, while others correspond to uncertain reasoning decisions. We present Token-Disentangled Latent Test-Time Scaling, an inference-time framework that makes latent refinement token-role-aware. Starting from an initial generated trajectory, we optimize a short hidden-state prefix while routing perception-side visual feedback to image-sensitive tokens and reasoning feedback to high-entropy tokens. Tokens selected by neither route are constrained by an anchor regularizer. Across both perception and reasoning benchmarks on Qwen2.5-VL-7B and InternVL3.5-8B, our method lifts macro accuracy over CoT by +2.57 and +1.51 respectively, and outperforms strong output-space test-time scaling baselines under matched decoded-candidate budgets. Code is available at https://github.com/Qwen-Applications/TD-LTTS.

Comments20 pages, 4 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑