arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.11315cs.AI

按推理需求路由:扩散视觉语言模型的轨迹感知解码控制

Routing by Reasoning Need: Trajectory-Aware Decoding Control for Diffusion Vision-Language Models

Yixiang Liu, Zhongxing Xu, Zhonghua Wang, Xiaoying Tang

首次发表
浏览论文内容

中文总结 AI 辅助

针对扩散视觉语言模型统一解码长度与推理需求不匹配的问题,提出无训练轨迹感知路由控制,按答案封闭性等信号动态调整解码策略,提升鲁棒性。

中文摘要 AI 辅助

扩散视觉语言模型通过迭代细化生成答案,在推理时暴露了可被检查和控制的中间答案轨迹。然而,这种可控性造成了推理需求不匹配的问题,即对具有不同推理需求的问题应用统一的生成长度。视觉上封闭的问题可能在稳定答案形成后因继续细化而受损,而对推理敏感的问题则可能因过早承诺而受损。我们将此问题表述为推理预算不匹配,并在LLaDA-V中对其进行研究。我们的无训练控制器不是选择统一的生成长度,而是利用来自答案封闭性、承诺证据和表示修订压力的轨迹信号,将每个示例路由到早期承诺、基线保留或推理支持解码,且不使用真实答案。在面向答案、混合推理和CoT敏感基准测试中,路由控制相比固定长解码、纯短解码和单规则干预提高了鲁棒性。这些收益不能仅用更短的输出来解释。答案封闭的示例通常受益于承诺,而CoT敏感的示例则需要保留或支持中间推理。综合来看,这些结果表明扩散VLM解码应根据观察到的轨迹所指示的状态来路由推理时控制,而不是依赖统一的解码长度。

英文摘要

Diffusion vision-language models generate answers through iterative refinement, exposing intermediate answer trajectories that can be inspected and controlled at inference time. However, this controllability creates a reasoning-need mismatch, where a universal generation length is applied to questions with different reasoning demands. Visually closed questions may be harmed by continued refinement after a stable answer has formed, whereas reasoning-sensitive questions may be harmed by premature commitment. We formulate this problem as reasoning-budget mismatch and study it in LLaDA-V. Rather than choosing a universal generation length, our training-free controller routes each example to early commitment, baseline preservation, or reasoning-supportive decoding using trajectory signals from answer closure, commitment evidence, and representation revision pressure, without using ground-truth answers. Across answer-focused, mixed-reasoning, and CoT-sensitive benchmarks, routed control improves robustness over fixed long decoding, pure short decoding, and single-rule interventions. The gains are not explained by shorter outputs alone. Answer-closed examples often benefit from commitment, whereas CoT-sensitive examples require preserving or supporting intermediate reasoning. Taken together, these results suggest diffusion VLM decoding should route inference-time control by the state suggested by the observed trajectory instead of relying on a universal decoding length.

发表机构

  • Southern University of Science and Technology(南方科技大学)
  • Monash University(莫纳什大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑