发表机构
Zhejiang University; The University of Hong Kong(浙江大学; 香港大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
EPOCH提出证据治理架构,通过任务契约、主动证伪和独立重放等机制治理搜索反馈,在AlgoTune等基准上超越基线,实现更可靠的科学发现。
AI 中文摘要
AI研究智能体越来越多地被用于搜索程序、数学构造和证明。然而,现有系统通常优化评估器反馈,却没有充分治理该反馈如何被解释、质疑和重用。因此,有前景但脆弱的候选可能被提升为发现,而基准改进、有限证书和定理级声明则太容易被混为一谈。我们引入EPOCH,一种旨在弥合这一差距的证据治理架构。EPOCH通过结合显式任务契约、类型化记忆、主动证伪、准入检查和独立重放来实现证据治理的发现循环,从而根据每个候选所支持的声明的强度和范围对其进行评估。EPOCH在AlgoTune上实现了最先进的总体性能,在平均归一化分数上大幅超过最强基线(0.65对比0.53),并在内部Math14套件上取得最高平均分数(0.57)。它在官方测试重放下展现出良好的保留行为,并在AgentHPO上引领描述性聚合。在十个发现问题上,EPOCH带来了实质性的任务特定进展,包括改进的可执行构造、优化算法、反例和证明支持的结果。这些进展展示了其将搜索转化为数学和计算领域具体进步的能力。总体而言,结果表明证据治理是迈向不仅产生更强解决方案、而且产生更可信科学发现的AI研究智能体的必要步骤。
英文摘要
AI research agents are increasingly used to search over programs, mathematical constructions, and proofs. However, existing systems typically optimize evaluator feedback without adequately governing how that feedback is interpreted, challenged, and reused. As a result, promising but fragile candidates can be promoted as discoveries, while benchmark improvements, finite certificates, and theorem-level claims are too easily conflated. We introduce EPOCH, an evidence-governed architecture designed to close this gap. EPOCH implements an evidence-governed discovery loop by combining explicit task contracts, typed memory, active falsification, admission checks, and independent replay, so that each candidate is evaluated against the strength and scope of the claim it supports. EPOCH achieves state-of-the-art aggregate performance on AlgoTune, substantially exceeding the strongest baseline in mean normalized score (0.65 vs. 0.53), and attains the highest mean score on the internal Math14 suite (0.57). It further shows favorable held-out behavior under official-test replay and leads the descriptive aggregate on AgentHPO. Across ten discovery problems, EPOCH delivers substantial task-specific advances, including improved executable constructions, optimized algorithms, counterexamples, and proof-supported results. These advances demonstrate its ability to convert search into concrete progress across mathematical and computational domains. Together, the results suggest that evidence governance is a necessary step toward AI research agents that produce not only stronger solutions, but also more trustworthy scientific discoveries.
Comments49 pages, 16 figures, including supplementary material