arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.37490cs.SC

CPUNeSy:控制模型写入以实现可靠的神经符号推理

CPUNeSy: Controlling Model Writes for Reliable Neuro-Symbolic Reasoning

Zeyan Li, Siyuan Qiu, Shuai Zhao, Jianfeng Xu

首次发表
浏览论文内容

中文总结 AI 辅助

CPUNeSy通过控制模型写入和证书门实现选择性弃权,结合确定性执行恢复多跳推理准确率,在多个基准上提升显著,但增益受基础器错误机制影响。

中文摘要 AI 辅助

大型语言模型(LLM)在回忆统计模式方面表现出色,但在需要推导答案时性能急剧下降,尤其是在多跳链上。将推导委托给确定性的符号执行器,将可靠性转移到模型生成的前提是否由源支持。我们引入了CPUNeSy,一种服务架构,通过任务定义的谓词接口和证书门控制模型对符号状态的写入,在基础(grounding)传递不一致时弃权(不执行)。组件分析隔离了确定性执行、受限基础、一致性和源复核。实验表明,在推导密集型任务上,确定性执行驱动了大部分准确率恢复;受控写入主要通过保留不支持或不一致的答案来提高选择性可靠性,但以覆盖率为代价。在法律和形式数学的多跳测试中,确定性执行恢复了对思维链和检索基线的大部分差距,全池增益高达35.0个百分点。认证是选择性服务控制,而非准确率机制:在ContractNLI上固定基础轨迹时,源复核移除了DeepSeek在两次投票一致性中幸存的错误答案的四分之一,但以可测量的覆盖率成本为代价。当弃权成本高昂时,将保留的案例路由到未认证的同模型回退,在MedCalc-Bench Verified上将Seed和DeepSeek的全池准确率分别提高了13.9和4.9个百分点;这些增益并非来自认证通道。在LeanDojo Benchmark 4上,内核受限池与BM25召回率@15(89.3%)相匹配。增益取决于基础器的错误机制:偏向偏差的基础器受益较少,这与我们的投票界限一致。证书保证相对于已接受前提的推导有效性;对自然语言源的语义保真度仍取决于源检查器,前瞻性验证是未来工作。

英文摘要

LLMs excel at recalling statistical patterns but degrade sharply when answers must be derived, especially on multi-hop chains. Delegating derivation to deterministic symbolic executors shifts reliability to whether model-generated premises are source-supported. We introduce CPUNeSy, a serving architecture that controls model writes to symbolic state via a task-defined predicate interface and certificate gate, abstaining when grounding passes disagree. Component analysis isolates deterministic execution, restricted grounding, agreement, and source rechecking. Experiments show deterministic execution drives most accuracy recovery on derivation-heavy tasks; controlled writes mainly improve selective reliability by withholding unsupported or inconsistent answers, at a coverage cost. On multi-hop tests in law and formal math, deterministic execution recovers most of the gap over chain-of-thought and retrieval baselines, with full-pool gains up to 35.0 points. Certification is selective-serving control, not accuracy mechanism: with grounding traces fixed on ContractNLI, source rechecking removes a quarter of DeepSeek's wrong answers surviving two-vote agreement, at measurable coverage cost. When abstention is costly, routing withheld cases to an uncertified same-model fallback raises full-pool accuracy on MedCalc-Bench Verified by 13.9 and 4.9 points for Seed and DeepSeek; these gains are not from the certified channel. On LeanDojo Benchmark 4, kernel-restricted pools match BM25 recall@15 (89.3%). Gains depend on the grounder's error regime: bias-dominated grounders benefit less, consistent with our voting bound. Certificates guarantee derivational validity relative to admitted premises; semantic faithfulness to natural-language sources remains conditional on the source checker, and prospective validation is future work.

发表机构

  • Shanghai Jiao Tong University(上海交通大学)

机构由 AI 辅助整理,请以论文原文为准。

↑