语义推理去噪:用语义算子修正语言模型的推理
Semantic Reasoning Denoising: Correcting Language Model Reasoning with Semantic Operators
浏览论文内容
中文总结 AI 辅助
本文提出语义推理去噪(SRD)方法,用语义算子修正语言模型推理,在六个域内基准平均提升基线3.2个点,七个跨数据集迁移目标上也有良好表现。
中文摘要 AI 辅助
大型语言模型可生成流畅的推理轨迹,但其局部语义错误会传播至错误结论,而无约束的自修正可能保留、放大或引入错误。现有扩散语言模型提供迭代优化,但通常将噪声定义为 token 掩码或替换,而非推理过程中的错误。本文提出语义推理去噪(Semantic Reasoning Denoising,SRD),一种针对自然语言推理轨迹的算子化马尔可夫去噪方法。SRD 用可执行的错误算子表示语义噪声,这些算子描述错误类型、位置以及被损坏和修复的命题。组合这些算子可构建逐渐更嘈杂的状态。训练期间,模型学习识别当前轨迹中活跃的语义噪声,并重建配对的相邻低噪声状态。推理期间,感知噪声水平的去噪会反复预测逆算子并检查其是否适用,因此每次执行的更新都会向稳定轨迹进行局部移动。在涵盖数学、代码、知识和常识的六个域内基准测试中,SRD 使最强的相同主干基线平均提升 3.2 个点。在七个跨数据集迁移目标上,它与 Llama-3-8B-Instruct 保持竞争力,并使最强的 Qwen3-8B 基线平均提升 2.9 个点。对噪声源、目标和去噪深度的分析进一步表明,结构化语义噪声预测和迭代算子执行是性能提升的核心。
英文摘要
Large language models can produce fluent reasoning traces whose local semantic errors propagate to an incorrect conclusion, while unconstrained self-correction may preserve, amplify, or introduce errors. Existing diffusion language models provide iterative refinement, but usually define noise as token masking or replacement rather than as errors in the reasoning process. We present Semantic Reasoning Denoising (SRD), an operatorized Markov denoising method for natural-language reasoning trajectories. SRD represents semantic noise with executable error operators that describe the error type, its location, and the corrupted and repaired propositions. Composing these operators constructs progressively noisier states. During training, the model learns to identify the semantic noise active in the current trajectory and to reconstruct the paired adjacent lower-noise state. During inference, noise-level-aware denoising repeatedly predicts an inverse operator and checks whether it is applicable, so each executed update makes a localized move toward a stable trajectory. Across six in-domain benchmarks spanning mathematics, code, knowledge, and commonsense, SRD improves the strongest same backbone baseline by 3.2 points on average. On seven cross-dataset transfer targets, it remains competitive with Llama-3-8B-Instruct and improves the strongest Qwen3-8B baseline average by 2.9 points. Analyses of noise sources, objectives, and denoising depth further show that structured semantic-noise prediction and iterative operator execution are central to the improvement.