arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Albilich:用于结合计算机代数系统(CAS)的基于大语言模型的数学研究的可操控证明状态编排工具

Albilich: Steerable Proof-State Orchestration for LLM-Based Mathematical Research with CAS Integration

Ting Gong, Michael Ruofan Zeng, Yong Yang

arXiv 2607.27705首次发表:更新:

AI 中文总结

Albilich是结合CAS等功能的开源数学自主研究智能体框架,在RealMath基准和Kourovka开放问题上取得优异效果,可实现AI辅助的可扩展数学研究。

AI 中文摘要

大型语言模型可为数学研究提供有用思路,但长期证明尝试仍难以协调、评估与复现。我们提出Albilich,这是一个用于数学自主研究的开源智能体框架,融合了长期推理、计算机代数系统(CAS)、文献检索及基于SQLite的持久上下文管理功能。我们在RealMath基准(Zhang等人,2025)和Kourovka笔记(Khukhro与Mazurov,2026)中的群论开放问题上对Albilich进行评估。它在RealMath基准上结合CAS时解决了10/10个问题,不使用CAS时解决了9/10个问题;在Kourovka问题上,Albilich针对问题21.142给出了反例,并对问题20.2的强化版本给出了证明。针对问题17.91的 ablation实验显示,启用CAS时令牌减少32.0%;针对问题21.142的 ablation实验显示,在缺少顾问智能体时,验证器拒绝率更高且无法合成证明路径。这些结果表明,Albilich是一个可人工操控、由CAS增强的环境,适用于可扩展的AI辅助数学研究。

英文摘要

Large language models can contribute useful ideas to mathematical research, yet long-horizon proof attempts remain difficult to coordinate, evaluate, and reproduce. We present Albilich, an open-source agentic harness for autoresearch in mathematics that combines long-horizon reasoning, computer algebra systems (CAS), literature retrieval, and persistent SQLite-based context management. We evaluate Albilich on the RealMath benchmark (Zhang et al. 2025) and on open problems in group theory from the Kourovka Notebook (Khukhro and Mazurov 2026). It solved 10/10 problems on RealMath with CAS and 9/10 with no CAS. On the Kourovka problems, Albilich produced a counterexample to Problem 21.142 and a proof of a strengthening of Problem20.2. Anablation on Problem 17.91 demonstrates 32.0% token reduction when CAS is enabled. An ablation on Problem 21.142 demonstrates higher verifier-rejection rate and failure to synthesize proof routes in the absence of the advisor agent. These results support Albilich as a human-steerable, CAS-boosted environment for scalable AI-assisted mathematical research.

Comments7 pages, comments welcome!

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑