arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.15278cs.CV

用于多步视觉推理的分层去噪

Hierarchical Denoising For Multi-Step Visual Reasoning

Zezhong Qian, Xiaowei Chi, Chak-Wing Mak, Tianze Zhou, Ruibin Yuan, Yuhan Rui, Hengzhe Sun, Zhuoqun Wu, Yuming Li, Siyuan Qian, Sirui Han, Shanghang Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对视频模型多步推理不足的问题,提出HDR框架,通过树形层次结构和稀疏分层注意力模式进行多步推理,在新基准测试中提升了成功率和进度,推理更快,数据效率更高,在机器人实验中展现潜力。

中文摘要 AI 辅助

视频模型正在演变为视觉基础模型,但仍缺乏类似人类的多步推理能力。流式自回归扩散模型高效但推理有限,双向扩散虽能全局修正但推理成本高。我们提出了HDR,一个将分层潜在因素集成到因果视频生成中进行多步推理的统一框架。HDR将视频潜在因素组织成树形层次结构,在流式输出前实现从粗到细的推理。粗去噪层保留不确定假设用于全局规划,细层逐步将其细化为具体视觉状态。稀疏分层注意力模式进一步降低时间注意力成本。我们引入了一个具有分布外情况的分层多步视频推理基准,涵盖六个任务。与流式自回归扩散基线相比,HDR将成功率从34.22提高到60.29,平均进度从76.00提高到89.56,推理速度比双向扩散快54.2倍,在仅2%训练数据时保留82.9%的全数据性能。真实世界机器人实验进一步证明了HDR在物理交互和世界建模方面的潜力。

英文摘要

Video models are evolving into vision foundation models, yet they still lack human-like multi-step reasoning. Streaming autoregressive diffusion models are efficient but limited in reasoning, while bidirectional diffusion enables global revision with high inference costs due to dense frame-level denoising. Both paradigms struggle to achieve logical consistency and low-latency streaming for complex reasoning tasks. We propose HDR (Hierarchical Denoising for Visual Reasoning), a unified framework that integrates hierarchical latents into causal video generation for multi-step reasoning. HDR organizes video latents into a tree-structured hierarchy, enabling coarse-to-fine reasoning before streaming output. Coarse denoising layers preserve uncertain hypotheses for global planning, while finer layers progressively refine them into concrete visual states. A sparse hierarchical attention pattern (SHAP) further reduces temporal attention costs. We introduce a level-stratified multi-step video reasoning benchmark with out-of-distribution cases, covering six tasks: maze navigation, Tower of Hanoi, one-line drawing, sliding puzzle, Sokoban, and water pouring. Compared with streaming autoregressive diffusion baselines, HDR improves success from 34.22 to 60.29 (76.2% relative gain) and increases average progress from 76.00 to 89.56, demonstrating more consistent reasoning trajectories. HDR maintains low-latency streaming at 0.70 seconds per latent, achieving 54.2 times faster inference than bidirectional diffusion. It also retains 82.9% of full-data performance with only 2% training data, compared with 52.0% for bidirectional diffusion. Real-world robot experiments further demonstrate HDR's potential for physical interaction and world modeling. Project demo: https://hierarchical-diffusion-reasoning.github.io/.

发表机构

  • State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(北京大学计算机科学学院多媒体信息处理技术国家重点实验室)
  • The Hong Kong University of Science and Technology(香港科技大学)
  • Beihang University(北京航空航天大学)
  • Fuzhou University(福州大学)
  • Muka Robotics(木卡机器人)

机构由 AI 辅助整理,请以论文原文为准。

↑