发表机构
University of the Chinese Academy of Sciences; Tencent(中国科学院大学; 腾讯)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对视觉潜在推理中连续潜在变量推理内容坍缩与离散潜在变量答案预测绕行的问题,提出CALR,通过功能锚定连接潜在形成与答案使用,结合潜在介导监督和推导级语义锚定,在五个数学基准上平均提升显著,较可比方法提升26.0个百分点。
AI 中文摘要
视觉潜在推理将渲染的推导过程压缩为紧凑的中间状态,从而减少文本推理的开销。现有方法在这些状态的表示方式上有所不同:连续方法避免了词汇表约束,而离散方法通过量化到有限码本提高了准确性。我们对具有代表性的连续和离散系统进行分析,确定了两个功能要求:答案必须依赖于潜在状态,且这些状态必须携带有效且针对问题的推理。连续潜在变量尽管推理内容坍缩,仍能影响答案;而离散潜在变量保留了可恢复的中间推理,但答案预测在很大程度上绕过了这些推理。为解决这些挑战,我们提出了连续锚定潜在推理(CALR),通过功能锚定将潜在状态的形成与答案的使用联系起来。借助来自信息平衡压缩的参考潜在变量,CALR将潜在介导的答案监督与推导级语义锚定相结合:前者将答案监督路由通过中间状态,后者将其解码内容锚定在针对问题的推导中。一种从并行到自回归的课程学习通过将后续潜在块的条件建立在已生成前缀之上,逐步发展序列推理。在跨模型家族的五个数学推理基准上的评估显示,准确率显著提升。在匹配的预算下,CALR相较于一种可比的连续潜在推理方法提升了26.0个百分点。进一步分析表明,其潜在变量支持答案预测并携带针对问题的中间推理。
英文摘要
Visual latent reasoning compresses rendered derivations into compact intermediate states, reducing textual reasoning overhead. Existing approaches differ in how they represent these states: continuous methods avoid vocabulary constraints, whereas discrete methods improve accuracy through quantization into a finite codebook. Our analysis of representative continuous and discrete systems identifies two functional requirements: answers must rely on latent states, and those states must carry valid, problem-specific reasoning. Continuous latents influence answers despite collapsed reasoning content, whereas discrete latents retain recoverable intermediate reasoning that answer prediction largely bypasses. To address these challenges, we propose Continuous Anchored Latent Reasoning (CALR), which connects latent formation with answer use through functional anchoring. With reference latents from information-balanced compression, CALR couples latent-mediated answer supervision with derivation-level semantic anchoring: the former routes answer supervision through intermediate states, while the latter grounds their decoded content in problem-specific derivations. A parallel-to-autoregressive curriculum develops sequential reasoning by conditioning subsequent latent blocks on generated prefixes. Evaluations on five mathematical reasoning benchmarks across model families show substantial accuracy gains. Under matched budgets, CALR gains 26.0 percentage points over a comparable continuous latent reasoning method. Further analyses show that its latents support answer prediction and carry problem-specific intermediate reasoning.