arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从错误到证明:最小核心引导的神经符号约束求解修复

From Errors to Proofs: Minimal-Core-Guided Repair for Neuro-Symbolic Constraint Solving

Dipankar Sarkar

arXiv 2608.14771首次发表:更新:

AI 中文总结

该研究提出最小核心引导的修复方法,将语言模型生成的不可满足程序的错误消息替换为最小不可满足核心,在77个问题的基准上使弱模型编造不可行问题解的比例从79%降至7%,核心价值是提供证书并拒绝编造解。

AI 中文摘要

让语言模型可靠求解约束问题,常需将问题转化为形式化规约,并将搜索任务委托给可靠的求解器。但转化本身是语言模型任务,不忠实的转化会导致求解器正确求解了错误问题。现有管道仅修复崩溃的转化,当程序运行但结果错误时,仅返回求解器的错误消息后便不再处理。我们用证明取代错误消息:当生成的程序不可满足时,我们从模型自身约束中提取最小不可满足核心,返回无法同时成立的精确集合,这是一种无泄漏的信号,可定位故障。在含精确预言的77个问题的新基准上,转化为Answer Set Programming(回答集编程)在7个领域中的6个是忠实的,仅在聚合覆盖调度上失败,该领域将转化负担集中在一种可诊断的模式中。与单纯错误不同,最小核心可阻止较弱模型为不可行问题编造解决方案,将编造比例从79%降至7%。而强思维链基线在准确率上与符号路线相当,因此该路线的价值不在于准确率,而在于其证书及拒绝编造的特性。

英文摘要

Making language models solve constraint problems reliably often means having them translate the problem into a formal specification and delegating the search to a sound solver. But the translation is itself a language-model task, and an unfaithful translation makes the solver faithfully solve the wrong problem. Existing pipelines repair only translations that crash, returning the solver's error message and falling silent when the program runs but is wrong. We replace the error message with a proof: when the generated program is unsatisfiable, we extract a minimal unsatisfiable core over the model's own constraints and hand it back the exact set that cannot hold together, a leakage-free signal that localizes the fault. On a new benchmark of 77 problems with an exact oracle, translation to Answer Set Programming is faithful on six of seven domains and fails only on aggregate coverage scheduling, which concentrates the translation tax in one diagnosable pattern. A minimal core, rather than a bare error, is what stops a weaker model from fabricating solutions to infeasible problems, cutting fabrication from 79% to 7%. A strong chain-of-thought baseline meanwhile matches the symbolic route on accuracy, so the route's value is not accuracy but certificates and its refusal to fabricate.

Comments7 pages, 2 figures. Accepted at the IJCAI-ECAI 2026 Workshop on Logic and Symbolic Reasoning (LogiSymb), poster

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑