arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一个合理的神经推理器剖析:在富含线索的填数问题中的一次性摊销、首次通过中毒和搜索惰性

Anatomy of a Sound Neural Reasoner: One-Shot Amortization, First-Pass Poisoning, and Search Inertness in Clue-Rich Completion

Aleksey Komissarov

arXiv 2607.19635首次发表:更新:

发表机构

Neapolis University, Pafos(帕福斯新帕福斯大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究富含线索填数问题中神经推理器的表现,发现LDT存在首次通过中毒等问题,添加CoLT可减少无效推导,还给出两种干预措施提升准确率,指出类似LDT系统是一次性摊销预测器,搜索主要消除计算浪费。

AI 中文摘要

神经求解器旨在推导、分支和修正中间状态。格点推导变换器(LDT)似乎也是如此。但在富含线索的数独中并非如此:一次前向传递就基本确定了整个网格(标准6x6上的每个空白单元格,增强9x9上的94 - 96%),将迭代求解器变成了一个包裹在精确验证器中的一次性预测器。所有硬切片失败在搜索开始前就已确定,首次通过时就自信地删除了正确解所需的值,我们称之为首次通过中毒。添加学习到的分支、MRV、回溯、值排除和共享冲突集(CoLT)并不会改变能解决的数独实例;它将重复的无效推导减少了1497倍。在固定训练预算下,仅约束图注意力就能达到全CoLT的准确率,而位置表只有在训练时间大幅延长时才能恢复,这表明了优化和样本效率优势而非绝对能力差异。诊断预测了两种有效干预措施。数字排列增强将9x9的准确率从低于1%提高到96.5 +/- 0.3(在对称不相交分割的三个训练种子上)。测试时对对称变换后的传递进行并集操作,在不重新训练的情况下将所有三个硬切片检查点从72.8 - 78.9%提高到100%。在从头开始的图着色中,一次性行为消失且搜索改变了准确率。在富含线索的填数问题中,类似LDT的系统是一次性摊销预测器而非学习到的搜索过程:准确率由校准和对称性决定while搜索主要是消除计算浪费。

英文摘要

Neural solvers are built to deduce, branch, and revise intermediate states. The Lattice Deduction Transformer (LDT) appears to do exactly that. In clue-rich Sudoku, it does not: one forward pass commits essentially the entire grid (every blank cell on standard 6x6, 94-96% on augmented 9x9), turning the iterative solver into a one-shot predictor wrapped in an exact verifier. All hard-slice failures are decided before search begins, when the first pass confidently deletes a value required by the true solution. We call this first-pass poisoning. Adding learned branching, MRV, backtracking, value exclusion, and shared nogoods (CoLT) does not change which Sudoku instances are solved; it cuts repeated invalid derivations 1,497-fold. At the frozen training budget, constraint-graph attention alone matches full-CoLT accuracy, while positional tables recover only under substantially longer training, indicating an optimization and sample-efficiency advantage rather than an absolute capacity difference. The diagnosis predicts two effective interventions. Digit-permutation augmentation raises 9x9 accuracy from below 1% to 96.5 +/- 0.3 across three training seeds on a symmetry-disjoint split. Test-time union over symmetry-transformed passes raises all three hard-slice checkpoints from 72.8-78.9% to 100% without retraining. On from-scratch graph coloring, one-shot behavior disappears and search changes accuracy. In clue-rich completion, LDT-like systems are one-shot amortized predictors rather than learned search procedures: accuracy is determined by calibration and symmetry, while search primarily removes computational waste.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑