arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Praxist:从实验产物到解决方案谱系

Praxist: From Experimental Artifacts to Solution Lineages

Jin Li, Ahmed Murtadha, Zhiyu Wang, Qiwen Chen, William Chen, Yifei Wu, Guan Wang, Andy L. Siy, Jiayi Yang, Mengsha Huang, Wenhao Li, Yixuan Liu, Shuailin Pan, Mingli Yuan, Sen Song, Yuhao Sun

arXiv 2608.25955首次发表:更新:

AI 中文总结

Praxist是一种以谱系为中心的代际系统,可将可复现产物与评估结果转换为证据图,在MLE-bench测试中成本仅为Claude Code基线的约十二分之一,同时在多类工程问题上实现性能提升。

AI 中文摘要

自主研发智能体如今可在自动化评估下编写、运行并改进可执行产物,但它们大多仅作为实验室工具:仅在精心挑选的基准测试中展示,其提升难以追溯到具体原因,且成本远超持续工程实践所能承受的范围。这一限制是结构性的:多数系统将每次尝试视为几乎独立的单元,因此日志、记忆和搜索树仅记录发生了什么,未明确是哪个设计元素带来了提升、其证据是否通过验证,或如何与其他元素重组。因此,长期研发过程会不断重复学习相同的经验教训。我们引入Praxist,这是一个以谱系为中心的代际系统,可将可复现产物和评估器结果转换为包含发现内容的类型化证据图、按路径划分的前沿领域及议程。将本地产物构建与群体层面的证据合成分离,使后续尝试可继承已验证的机制、未解决的主张及有用约束,并让结果与可检查的谱系关联。在标准化的75任务MLE-bench套件上,官方评分器的最终结果显示,Praxist获得60枚奖牌(占80.0%),其中49枚为金牌;而Claude Code基线模型在Claude Opus 4.8上获得55枚奖牌(占73.3%),其中34枚为金牌,记录的模型成本分别为3054美元和38370美元,Praxist的成本约为其十二分之一。四个案例研究——量化交易、激光雷达-惯性-视觉SLAM、托卡马克磁控制及火箭着陆——将相同流程应用于开放式工程问题,在每项任务的原生基线的核心准确率、存续性或资源成本上均有提升,且发现路径已记录在案。据我们所知,本文首次汇集了成本低一个数量级的更强产物,每个产物都有可审计的谱系,这是实际生产研究所需的操作模式,而非基准测试展示的操作模式。

英文摘要

Autonomous R\&D agents now write, run, and improve executable artifacts under automated evaluation---but largely as laboratory instruments: shown on curated benchmarks, with gains that are hard to trace to a cause and costs well above what sustained engineering practice absorbs. The limitation is structural. Most systems treat each attempt as nearly self-contained, so logs, memories, and search trees record what happened without establishing which design element produced an improvement, whether its evidence survived validation, or how it recombines with others. Long campaigns therefore keep re-learning the same lessons. We introduce Praxist, a lineage-centered generational system that converts reproducible artifacts and evaluator outcomes into a typed evidence graph of findings, lane-structured frontiers, and agendas. Separating local artifact construction from cohort-level evidence synthesis lets later attempts inherit validated mechanisms, unresolved claims, and useful constraints, and leaves results attached to an inspectable lineage. On the standardized 75-task MLE-bench suite, the finalized official-grader results give Praxist 60 medals (80.0\%), 49 of them gold, against 55 medals (73.3\%) and 34 gold for a Claude Code baseline on Claude Opus 4.8---at a recorded model spend of US\$3,054 versus US\$38,370, roughly a twelfth of the cost. Four case studies---quantitative trading, LiDAR-inertial-visual SLAM, tokamak magnetic control, and rocket landing---carry the same process into open-ended engineering problems, improving on each task-native baseline in headline accuracy, survival, or resource cost, with the discovery path on record. Stronger artifacts at an order of magnitude less spend, each backed by an auditable lineage, are, to our knowledge, first brought together here: the operating profile production research requires, not the one a benchmark demonstration establishes.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑