arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

思考稀疏,预测密集:用于图像超分辨率的连续思维机器

Think Sparse, Predict Dense: Continuous Thought Machines for Image Super-Resolution

Zekai Shi

arXiv 2607.18856首次发表:更新:

发表机构

Xi’an Jiaotong University(西安交通大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对图像超分辨率问题,提出窗口级连续思维机器(CTM)及DQ-CTM机制,将紧凑思维表示转换为密集查询,通过ThinkSR实例进行实验,结果显示PSNR等指标提升,证明稀疏潜在思维用于密集空间重建可行,推动相关架构发展。

AI 中文摘要

连续思维机器引入了一个内部时间维度,其中神经元级历史和同步派生表示在一系列思维滴答中演变。将此机制扩展到密集视觉预测并非易事,因为图像超分辨率等任务需要在每个输出位置都有空间证据,而不是压缩成单个全局表示。在提出的窗口级CTM使用中,思维动态为每个局部窗口生成一个紧凑的摘要表示。DQ-CTM通过结构化低秩、参数高效的紧凑到密集查询机制将此紧凑思维表示转换为窗口对齐的密集查询。窗口内的每个位置都接收自己的查询,而共享的思维动态在滴答间逐步细化密集表示。在其超分辨率实例ThinkSR中,编码特征图在无令牌池的情况下被划分为局部视觉窗口,在共享细化后恢复到原始特征域,并解码为高分辨率图像。在固定的四个滴答训练范围内的初步实验揭示了渐进的重建轨迹。PSNR-Y从T=0时的28.1045dB增加到T=4时的30.2817dB,PSNR-RGB从26.6271dB增加到28.7781dB,平均l1误差从0.034602降低到0.023545。所有100张评估图像从T=1到T=4都有改善。这些初步结果确立了稀疏潜在思维用于密集空间重建的可行性,并激发了更广泛的用于密集视觉的连续思维架构。

英文摘要

Continuous Thought Machines introduce an internal temporal dimension in which neuron-level histories and synchronization-derived representations evolve over a sequence of thought ticks. Extending this mechanism to dense visual prediction is non-trivial, because tasks such as image super-resolution require spatial evidence to remain available at every output location rather than being compressed into a single global representation. In the proposed window-level use of CTM, the thought dynamics produce a compact summary representation for each local window. DQ-CTM transforms this compact thought representation into window-aligned dense queries through a structured low-rank, parameter-efficient compact-to-dense query mechanism. Each position within a window receives its own query, while shared thought dynamics progressively refine the dense representation across ticks. In its super-resolution instantiation, termed ThinkSR, encoded feature maps are partitioned into local visual windows without token pooling, restored to the original feature field after shared refinement, and decoded into a high-resolution image. Preliminary experiments under a fixed four-tick training horizon reveal a progressive reconstruction trajectory. PSNR-Y increases from 28.1045 dB at $T=0$ to 30.2817 dB at $T=4$, while PSNR-RGB increases from 26.6271 dB to 28.7781 dB and the mean $\ell_1$ error decreases from 0.034602 to 0.023545. All 100 evaluated images improve from $T=1$ to $T=4$. These initial results establish the feasibility of sparse latent thought for dense spatial reconstruction and motivate broader continuous-thought architectures for dense vision.

Comments8 pages, 6 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑