发表机构
School of Computer Science, Wuhan University; Artificial Intelligence Innovation and Incubation, Fudan University; Department of Computer Science, Faculty of Science, University of Bath; College of Computer Science and Artificial Intelligence, Fudan University(武汉大学计算机科学学院; 复旦大学人工智能创新与孵化; 巴斯大学理学院计算机科学系; 复旦大学计算机科学与人工智能学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对科学推理图提取中LLMs输出有问题的情况,提出无需训练的PEARL框架,通过特定模式和反馈修复推理图。在ARCHE基准测试中,该框架提升了严格通过率和平均REA,为相关工作流程提供可靠性层。
AI 中文摘要
科学推理图提取(SRGE)旨在恢复观察、证据、中间主张和论文级结论之间的明确联系。大型语言模型(LLMs)能生成类似图的科学解释,但输出常存在语法错误、边缘标签漂移、根节点方向错误和源锚点薄弱等问题。我们提出了PEARL(通过抽象和修复层的皮尔士提取),这是一个无需训练的框架,能将有噪声的LLM图响应转化为可审计的推理图,并将其修复至严格的语义有效性。PEARL首先在封闭的皮尔士模式下实现明确的图内容,然后使用匹配的基于证据的判断反馈来修复被拒绝的边缘类型、局部推理步骤和终端根节点,同时保留审计跟踪。在来自ARCHE(一个潜在推理链提取基准)的五个包含70篇论文的模型档案上,PEARL将LLM基线的严格通过率从0/350提高到300/350,平均REA从0.339提高到0.906。这些图为需要可检查推理痕迹而非无约束图生成的研究代理和人工智能科学家工作流程提供了一个可靠性层。代码和审计工件可在该https URL获取。
英文摘要
Scientific Reasoning Graph Extraction (SRGE) aims to recover explicit links among observations, evidence, intermediate claims, and paper-level conclusions. LLMs can produce graph-like scientific explanations, but their outputs often mix malformed syntax, drifting edge labels, incorrectly oriented roots, and weak source anchors. We propose PEARL (Peircean Extraction via Abstraction and Repair Layer), a training-free framework that turns noisy LLM graph responses into auditable reasoning graphs and repairs them toward strict semantic validity. PEARL first materializes explicit graph content under a closed Peircean schema, then uses matched evidence-grounded judge feedback to repair rejected edge types, local inference steps, and terminal roots while preserving an audit trail. On five 70-paper model archives from ARCHE, a benchmark for latent reasoning-chain extraction, PEARL raises strict gate passes from 0/350 for the LLM baseline to 300/350, with average REA improving from 0.339 to 0.906. The graphs provide a reliability layer for research-agent and AI scientist workflows that need inspectable reasoning traces rather than unconstrained graph regeneration. Code and audit artifacts are available at https://github.com/BohanSu/auditable-repair-reasoning-graphs/tree/300-350_workshop .
CommentsAccepted at WAICA 2026 Multi-Modal Agents for Science Workshop