arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

约束锚定推理轨迹

Constraint-Anchored Reasoning Traces

Zehua Cheng, Wei Dai, Jiahao Sun

arXiv 2607.16727首次发表:更新:

发表机构

University of Oxford(牛津大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究自回归多模态大语言模型错误滚雪球问题,提出神经符号框架CART,将自然语言推理与符号约束断言交织,经双管齐下的模块验证约束,防止错误传播,在多基准测试中降低滚雪球率、提高准确率并控制推理开销。

AI 中文摘要

自回归多模态大语言模型存在错误滚雪球问题,即思维链早期的单个错误推理会使所有下游推理出错。在65%的此类情况下,一旦出现第一个错误,推理就会在所有剩余步骤中失败。现有缓解方法存在缺乏符号基础、纠错太晚或牺牲自然语言推理灵活性等问题。我们提出约束锚定推理轨迹(CART),这是一种神经符号框架,训练大语言模型将自然语言推理步骤与符号约束断言交织。通过结合学习到的神经基础头和布尔约束传播的双管齐下的约束传播模块,持续验证这些锚点并检查逻辑一致性。检测到矛盾时,回溯控制器停止生成并恢复到最后一个一致的检查点,防止错误传播。可变频率发射机制允许模型自适应控制锚点密度,避免轨迹膨胀。我们通过增强GQA、CLEVR - CoGenT和VCR构建了218K训练实例,并通过LoRA微调开源大语言模型。在五个基准测试中,CART将滚雪球率从0.65降至0.14,提高了GQA准确率4.6个百分点,在POPE上达到89.1 F1,推理开销最多为18%。

英文摘要

Autoregressive multimodal large language models (MLLMs) suffer from error snowballing: a single incorrect inference early in a chainof-thought (CoT) trace corrupts all downstream reasoning. We find that in state-of-the-art open-source MLLMs, once the first error occurs, the reasoning cascades into failure across all remaining steps in 65% of such cases (a metric we term the snowball rate). Existing mitigations-sampling multiple chains, post-hoc self-verification, or full program synthesis-either lack symbolic grounding, catch errors too late, or sacrifice the flexibility of natural language reasoning. We propose Constraint-Anchored Reasoning Traces (CART), a neuro-symbolic framework that trains MLLMs to interleave natural language reasoning steps with symbolic constraint assertions: lightweight, machine-checkable statements about visual content (e.g., count(red_objects) = 3). A dual-pronged Constraint Propagation Module-combining a learned neural grounding head with Boolean Constraint Propagation-continuously verifies these anchors against extracted visual features and checks their mutual logical consistency. When a contradiction is detected, a backtrack controller halts generation and reverts to the last consistent checkpoint, preventing error propagation. A variable-frequency emission mechanism allows the model to adaptively control anchor density, avoiding trace bloat. We construct 218K training instances by augmenting GQA, CLEVR-CoGenT, and VCR with ground-truth constraint annotations derived from scene graphs, and fine-tune open-source MLLMs (LLaVA-NeXT, Qwen2-VL) via LoRA. On five benchmarks, CART reduces the snowball rate from 0.65 to 0.14, improves GQA accuracy by +4.6 percentage points over trainingonly baselines, and achieves 89.1 F1 on POPE-all with at most 18% inference overhead.

Comments15 pages, ACM MM 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑