因果锻造:一个用于因果推理自动化研究的形式化基础、自我改进的智能框架
CausalSmith: A Formally Grounded, Self-Improving Agentic Framework for Automated Research in Causal Inference
浏览论文内容
中文总结 AI 辅助
提出因果锻造框架用于因果推理自动化研究,结合因果论库与因果史密斯智能管道,通过声明审核增强内核验证,并用自主研究工件评估系统,为因果推理自动化研究提供新途径。
中文摘要 AI 辅助
自动化理论研究不仅受候选结果生成的限制,还受其可靠评估的制约。常见方法是用大语言模型(LLM)审阅者来闭合研究循环,但此类审阅者在经验上不可靠。我们提出因果锻造,一个基于精益证明助手的因果推理自动化理论研究框架。它结合了因果论(一个包含7035个机器检查声明的因果推理基础精益库)和因果史密斯(一个自我改进的智能管道)。该管道通过声明审核增强内核验证,将每个形式定理与其要表达的非正式声明进行比较。我们用自主研究运行产生的工件评估了该系统。
英文摘要
Automating theoretical research requires generating candidate results and evaluating them reliably. Models keep getting better at the first, while the second remains hard. A common approach asks one large language model (LLM) to review what another produced, yet such reviewers are empirically unreliable: they may accept fabricated papers and catch the fabrication at close to chance rates~\citep{badscientist2025}. We present \textsc{CausalSmith}, a framework for automated theoretical research in causal inference built on the Lean proof assistant, where a proof is checked by a program rather than read by a referee. \textsc{CausalSmith} rests on \textsc{Causalean}, a foundational Lean library for causal inference holding 8,179 machine-checked definitions and theorems, developed with language-model assistance under human design and review. Around it, we build a self-improving agentic pipeline that selects research topics, proposes results, formalizes statements, constructs proofs, and presents the resulting artifacts for human inspection. Moreover, the pipeline pairs Lean verification with a statement audit that compares each formal theorem against the informal claim behind it. We evaluate the system using artifacts produced by completed autonomous research runs. The source code, formal library, and run records are available at https://github.com/Jiyuan-Tan/CausalSmith.