arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.01534cs.CV

用于通用符号推理的递归视觉语言模型

Recursive Vision Language Models for General Symbolic Reasoning

Omid Nejati Manzari, Guillaume Lajoie, Hassan Rivaz

首次发表
浏览论文内容

中文总结 AI 辅助

提出基于预训练Qwen的递归推理框架R-Qwen,结合递归模型与预训练LLMs的优势,在8个基准测试中优于同类模型和更大LLMs,ARC-AGI上提升27.6%。

中文摘要 AI 辅助

数独、迷宫寻路和ARC等困难符号推理任务对大型语言模型(LLMs)而言仍具挑战性,因为LLMs固定深度的自回归推理限制了系统性搜索、优化和回溯。尽管递归模型如分层推理模型(HRM)和微型递归模型(TRM)通过迭代隐态优化解决了这一局限,但它们通常是特定任务的,且未利用预训练语言先验。我们提出R-Qwen,这是一个基于预训练Qwen主干构建的递归推理框架。R-Qwen通过程序化自递归和深度监督反复优化候选解,将递归模型的结构化迭代计算与预训练LLMs的语言和推理先验相结合。我们进一步将分层监督加权(HSW)适配到自回归模型,对递归步骤的损失进行指数加权。HSW至少降低了50%的梯度方差,提高了随机梯度的信噪比,并加速了收敛。在8个具有挑战性的基准测试中,R-Qwen在使用可训练参数数量相当的情况下,始终优于现有的递归推理模型和规模大得多的LLMs。值得注意的是,在ARC-AGI数据集上,我们的模型比基线实现了27.6%的提升,凸显了递归优化对通用符号推理的有效性。这些结果表明,递归推理机制和预训练语言模型先验是改进符号谜题求解的互补方法。代码和模型将在录用后发布。

英文摘要

Hard symbolic-reasoning tasks such as Sudoku, maze pathfinding, and ARC remain challenging for LLMs due to their fixed-depth autoregressive reasoning, which limits systematic search, refinement, and backtracking. While recursive models such as Hierarchical Reasoning Model (HRM) and Tiny Recursive Model (TRM) address this limitation through iterative latent-state refinement, they are typically task-specific and do not leverage pretrained language priors. We propose R-Qwen, a recursive reasoning framework built upon a pretrained Qwen backbone. R-Qwen repeatedly refines a candidate solution through programmatic self-recursion and deep supervision, combining the structured iterative computation of recursive models with the linguistic and reasoning priors of pretrained LLMs. We further adapt Hierarchical Supervision Weighting (HSW) to autoregressive models by exponentially weighting losses across recursive steps. HSW reduces gradient variance by at least 50\%, improves the signal-to-noise ratio of stochastic gradients, and accelerates convergence. Across eight challenging benchmarks, R-Qwen consistently outperforms prior recursive reasoning models and substantially larger LLMs while using a comparable number of trainable parameters. Notably, on ARC-AGI dataset, our model achieves a 27.6\% improvement over the baseline, highlighting the effectiveness of recursive refinement for general symbolic reasoning. These results suggest that recursive reasoning mechanisms and pretrained language model priors are complementary approaches for improving symbolic puzzle-solving. Code and models will be released after acceptance.

↑