arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

推理捷径与价值对称性:对称性允许什么、架构实现什么以及优化选择什么

Reasoning Shortcuts and Value Symmetries: What Symmetry Permits, Architecture Realizes, and Optimization Selects

Xin Xu

arXiv 2608.10420首次发表:更新:

发表机构

Carnegie Mellon University; University of Pennsylvania(卡内基梅隆大学; 宾夕法尼亚大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究分析神经符号系统的推理捷径,发现现有框架的关键定义不适用于其评估的基准,重新测量揭示无法解释的对比例与可证明结构相关,还明确了对称性惰性、自同构存在性等问题的计算复杂度,区分了对称性允许与优化选择的内容。

AI 中文摘要

推理捷径是神经符号系统规则的解决方案,这些方案通过非预期概念生成正确预测。Takemura、Inoue和Nishino近期提出的框架通过价值重标记的自同构群对其进行分析,并提出核心开放问题:规则何时能确定概念?我们首先表明,该框架的关键定义——在每个位置应用相同的置换——按所述方式不适用于其评估的四个异构基准中的任何一个,而最直接的嵌入方式(将域填充为相同大小)会产生置信的虚假病理结果:在CLE4EVR上报告的90.91%的解决方案对无法解释,而我们引入的层次结构中每个明确定义的成员报告的比例为0%,且填充后的结果内容随配置文件排序而变化。在15个预定义预测(其中13个已确认)下重新测量11个规则族,无法解释的对的比例范围为0%至99.9999%,且与可证明的结构相关:6个定理给出传递性及其失效的充分条件,包括自由槽引理,可仅从语法上证明Kandinsky的病理情况。对于电路给定的规则,判断坐标的对称性惰性是coNP完全问题;非平凡自同构的存在性在随机归约下是coNP难问题,属于Σ₂^p,除非多项式层次结构(PH)崩溃,否则不是Σ₂^p完全问题,在单调电路上则直接是coNP完全问题。在布尔情况下,传递性被精确分类:当解集合是仿射陪集时,自同构可解释一切。弱监督模型将所有94个观测到的推理捷径置于分量理论标记的层级,而在其证明传递性的48个层级中无任何捷径;12个类型模糊层级未产生任何捷径,从而区分了对称性允许的内容与优化选择的内容,且双路控制(dual-head control)复制了该分布。所有数字均来自已发布的人工制品。

英文摘要

Reasoning shortcuts are rule solutions that reach correct predictions through unintended concepts. A recent framework of Takemura, Inoue, and Nishino analyzes them through an automorphism group of value relabelings, asking when rules pin concepts down. Its key definition, one value permutation shared across all positions, does not apply as stated to any of its four heterogeneous benchmarks, and the most direct embedding, padding, produces confident false pathology: 90.91% of solution pairs unexplained on CLE4EVR, versus 0% under every well-defined rung of the componentwise hierarchy we introduce; the padded verdict rotates under configuration-file ordering. Across eleven rule families under fifteen pre-specified predictions (thirteen confirmed), unexplained-pair rates span 0% to 99.9999% and track provable structure: six theorems give sufficient conditions for transitivity and its failure. For circuit-given rules, symmetry-inertness of a coordinate is coNP-complete; automorphism existence is coNP-hard under randomized reductions, lies in $Σ_2^p$, is not $Σ_2^p$-complete in the Boolean case unless PH collapses, and is coNP-complete on monotone circuits. Boolean transitivity is classified exactly: automorphisms explain everything iff the solution set is an affine coset. Weakly supervised models place all 94 observed shortcuts at the one level the theory flags, none at the 48 it certifies transitive, and none at twelve typed-ambiguous levels. Relocating the absorbing element moves every shortcut with it; a confusion null attributes the location to geometry while the observed rate exceeds it by half again. Trained end to end on CLE4EVR's rule and heterogeneous domains through a synthetic prototype front end, models produce 20,223 label-preserving errors with zero different-orbit exceptions, as transitivity predicts, where the padded instrument would misreport 78-88% of them.

Comments62 pages, 2 figures, 8 tables. Under review

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑