arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2509.21193cs.CLcs.AI

Eigen-1:基于监控检索的自适应多智能体精炼用于科学推理

Eigen-1: Adaptive Multi-Agent Refinement with Monitor-Based RAG for Scientific Reasoning

  • Yale University(耶鲁大学)
  • Shanghai Jiao Tong University(上海交通大学)
  • Fudan University(复旦大学)
  • University of California, Los Angeles(加州大学洛杉矶分校)
  • Shanghai AI Lab(上海人工智能实验室)
  • University of Oxford(牛津大学)
  • Eigen AI

机构由 AI 辅助整理,请以论文原文为准。

Xiangru Tang, Wanghan Xu, Yujie Wang, Zijie Guo, Daniel Shao, Jiapeng Chen, Cixuan Zhang, Ziyi Wang, Lixin Zhang, Guancheng Wan, Wenlong Zhang, Lei Bai, Zhenfei… 展开作者

Xiangru Tang, Wanghan Xu, Yujie Wang, Zijie Guo, Daniel Shao, Jiapeng Chen, Cixuan Zhang, Ziyi Wang, Lixin Zhang, Guancheng Wan, Wenlong Zhang, Lei Bai, Zhenfei Yin, Philip Torr, Hanrui Wang, Di Jin

更新

AI总结:

提出Eigen-1框架,结合隐式检索与分层精炼,在HLE生物/化学上达48.3%准确率,显著优于基线,同时降低令牌和步骤开销。

AI中文摘要:

大型语言模型(LLM)近期在科学推理方面展现出强劲进展,但仍存在两大主要瓶颈。首先,显式检索割裂了推理过程,强加了额外的令牌和步骤的隐性“工具税”。其次,多智能体流水线往往通过对所有候选方案进行平均而稀释了强解。我们通过一个结合隐式检索和结构化协作的统一框架来解决这些挑战。在该框架的基础上,一个基于监控的检索模块在令牌级别运作,以最小干扰将外部知识整合到推理中。在此基础之上,层次化解精炼(HSR)迭代地将每个候选方案指定为锚点,由其同伴进行修复,而质量感知迭代推理(QAIR)则根据解的质量调整精炼过程。在“人类最后考试”(HLE)生物/化学金牌子集上,我们的框架达到了48.3%的准确率——这是迄今报告的最高水平,比最强的智能体基线高出13.4个百分点,比领先的前沿LLM高出最多18.1个百分点,同时将令牌使用量减少了53.5%,智能体步骤减少了43.7%。在SuperGPQA和TRQA上的结果证实了跨领域的稳健性。错误分析显示,推理失败和知识缺口在超过85%的案例中同时出现,而多样性分析揭示了一个明显的二分法:检索任务受益于解的多样性,而推理任务则倾向于共识。这些发现共同表明,隐式增强和结构化精炼克服了显式工具使用和统一聚合的低效性。代码可在以下网址获取:https://github.com/tangxiangru/Eigen-1。

英文摘要:

Large language models (LLMs) have recently shown strong progress on scientific reasoning, yet two major bottlenecks remain. First, explicit retrieval fragments reasoning, imposing a hidden "tool tax" of extra tokens and steps. Second, multi-agent pipelines often dilute strong solutions by averaging across all candidates. We address these challenges with a unified framework that combines implicit retrieval and structured collaboration. At its foundation, a Monitor-based retrieval module operates at the token level, integrating external knowledge with minimal disruption to reasoning. On top of this substrate, Hierarchical Solution Refinement (HSR) iteratively designates each candidate as an anchor to be repaired by its peers, while Quality-Aware Iterative Reasoning (QAIR) adapts refinement to solution quality. On Humanity's Last Exam (HLE) Bio/Chem Gold, our framework achieves 48.3\% accuracy -- the highest reported to date, surpassing the strongest agent baseline by 13.4 points and leading frontier LLMs by up to 18.1 points, while simultaneously reducing token usage by 53.5\% and agent steps by 43.7\%. Results on SuperGPQA and TRQA confirm robustness across domains. Error analysis shows that reasoning failures and knowledge gaps co-occur in over 85\% of cases, while diversity analysis reveals a clear dichotomy: retrieval tasks benefit from solution variety, whereas reasoning tasks favor consensus. Together, these findings demonstrate how implicit augmentation and structured refinement overcome the inefficiencies of explicit tool use and uniform aggregation. Code is available at: https://github.com/tangxiangru/Eigen-1.

↑