arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.11341cs.AI

Apodex Discovery:用于评估和构建发现式人工智能的现实基准与环境

Apodex Discovery: Reality Benchmarks and Environments for Evaluating and Building Discoverative Artificial Intelligence

Brian Wang, Bin Feng, Xiaoman Pan, Chenyang An, Felix Liu, Tangqi Fang, Gongbo Sun, Lingfeng Shen, Ning Wang, Handuo Zhang, Feng Chen, Fuchao Yang, Xiang Wang, … 展开作者

Brian Wang, Bin Feng, Xiaoman Pan, Chenyang An, Felix Liu, Tangqi Fang, Gongbo Sun, Lingfeng Shen, Ning Wang, Handuo Zhang, Feng Chen, Fuchao Yang, Xiang Wang, Jiacheng Lin, Siting Li, Zixuan Liu, Chi Han, Zhenhailong Wang, Kunlun Zhu, Lawrence Zhao, Yueqi Guo, Kailong Wen, Feng Xing, Yiling Guo, Lidong Bing, David Tan, Bo An, Heng Ji, Sheng Wang

AI总结:

Apodex Discovery是评估和构建发现式AI的框架,含三个核心组件,在AAV衣壳设计等任务中较现有方法取得性能提升,推动AI评估向可验证的真正发现研究发展。

AI中文摘要:

阿波罗登月计划的成功并非仅因工程师能求解复杂方程,更在于将遥远的目标转化为包含明确目标、模拟、验证及反复修正的任务架构。人工智能如今面临类似转变:前沿模型在问题、工具和成功标准明确时可解决复杂任务,但现实中重要挑战极少以可执行或可验证的形式呈现。我们推出Apodex Discovery,这是一个通过重型求解器构建和评估发现式人工智能的框架,该系统包含基础模型、工具链、工具及控制策略,用于开展扩展的、有状态的、可验证的研究。它有三个核心组件:第一,问题侦察流程调研了16个行业的561个产业,收集了423个高价值现实问题,并选出20个用于初始版本发布;第二,通用环境-任务-片段抽象提供数据、工具、约束、反馈、轨迹记录,以及中间产物和最终提交的验证;第三,HDS6独立于任务最终成功情况评估工具、修复、替代方案、连贯性、证据及范围。在AAV衣壳设计中,Apodex在活力、向性、结构预测及生成设计方面较已发表的最新水平提升了7%;在药物重定位与重新配方任务中,特定任务的生物医学环境使GPT-5.5和GPT-5.6-sol的平均归一化预测分数较相同闭卷主干模型分别提升了2.5和7.6个百分点。控制消融实验表明,固定的TRACES片段界面可将性能差异归因于特定求解器组件。Apodex Discovery将人工智能评估从预定义基准推进到旨在实现真正发现的可验证研究。

英文摘要:

Apollo did not reach the Moon merely because its engineers could solve difficult equations. It succeeded by turning a distant ambition into a mission architecture of explicit objectives, simulation, verification, and repeated correction. AI now faces a similar transition: frontier models can solve difficult tasks once the problem, tools, and success criteria are specified, yet consequential real-world challenges rarely arrive in an executable or verifiable form. We introduce Apodex Discovery, a framework for building and evaluating discoverative AI through the heavy-duty solver, a system comprising a foundation model, harness, tools, and control policies that pursues extended, stateful, verifiable investigations. It has three core components. First, a problem-scouting process surveyed 561 industries across 16 sectors, assembled 423 high-value real-world problems, and selected 20 for the initial release. Second, a common environment-task-episode abstraction provides data, tools, constraints, feedback, trajectory recording, and verification of intermediate artifacts and final submissions. Third, HDS6 evaluates Tools, Repair, Alternatives, Coherence, Evidence, and Scope independently of final-task success. In AAV capsid design, Apodex surpassed the published state of the art by 7% across viability, tropism, structure prediction, and generative design. In drug repurposing and reformulation, a task-specific biomedical environment improved the mean normalized prediction score of GPT-5.5 and GPT-5.6-sol by 2.5 and 7.6 points over the same closed-book backbone. Controlled ablations show that the fixed TRACES episode interface enables attribution of performance differences to specific solver components. Apodex Discovery moves AI evaluation beyond predefined benchmarks toward verifiable investigations aimed at genuine discovery.

补充信息

↑