arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.12036cs.AIcs.CLcs.HCcs.LGcs.MA

Mechanist:作为科学仪器的AI,用于发现智能的机制

Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence

  • Zhejiang University(浙江大学)
  • National University of Singapore(新加坡国立大学)
  • Heriot-Watt University(赫瑞瓦特大学)
  • Southern University of Science and Technology(南方科技大学)
  • University of California, San Diego(加利福尼亚大学圣迭戈分校)
  • Northeastern University(东北大学)

机构由 AI 辅助整理,请以论文原文为准。

Mengru Wang, Junfeng Fang, Shuofei Qiao, Zhenqian Xu, Haoming Xu, Haoxiong Wang, Shumin Deng, Linyi Yang, Xin Xu, Yunzhi Yao, Dan Zhang, Fei Shen, Zhixiang Cui,… 展开作者

Mengru Wang, Junfeng Fang, Shuofei Qiao, Zhenqian Xu, Haoming Xu, Haoxiong Wang, Shumin Deng, Linyi Yang, Xin Xu, Yunzhi Yao, Dan Zhang, Fei Shen, Zhixiang Cui, Buqiang Xu, Haozhe Luo, Yunxiang Wei, Ningyu Zhang, Julian McAuley, Tat Seng Chua, Huajun Chen

AI总结:

本研究推出智能体系统Mechanist,以AI为科学仪器自主发现AI智能机制,其生成假设更有价值、实验更可靠,还能发现安全风险、构建信念机制理论并转化为提升模型性能的干预措施。

AI中文摘要:

AI模型在多个领域取得了显著成功,但其能力背后的机制以及可能带来的风险仍未被充分理解。随着AI发展速度加快且日益自动化,机制探索仍主要依赖人工,这扩大了模型能力与我们理解、控制模型的能力之间的差距。为弥合这一差距,我们推出了Mechanist,这是一个将AI作为科学仪器、用于自主发现AI智能背后机制的智能体系统。为支持自主机制发现,我们构建了一个聚焦可解释性的知识图谱,包含约13000篇论文,并将其与涵盖26个领域的4300万篇论文的多学科数据库整合。我们还整理了一个包含32种基础方法的库,用于机制分析、因果干预和验证。与Claude Code及现有的AI科学家系统相比,Mechanist能生成更有价值的机制假设,且实验执行更可靠。Mechanist还展现出从发现模型行为到解释和控制AI模型的进展:首先,它在科学实验室中发现了一种反直觉的安全风险,即不安全特性可通过看似安全的训练数据跨模态传递;随后,它开发了一种信念的机制理论,揭示模型如何表征世界知识、形成信念、推断他人信念,以及这些机制在预训练过程中如何产生;最后,它将这些机制见解转化为实际干预措施,在多种场景下提升模型性能,并引导科学基础模型生成具有指定特性的DNA序列。

英文摘要:

AI models are increasingly used in scientific discovery and human decision-making. Yet how AI models work and what risks they pose remain poorly understood. As AI development becomes faster and more automated, research on the mechanisms underlying AI remains largely manual. To bridge this gap, we introduce Mechanist, an agentic system that uses AI as a scientific instrument for the autonomous discovery of mechanisms underlying AI. To ground novel mechanism hypotheses, we construct a scientific knowledge graph of 13,000 studies on AI mechanisms, alongside a multidisciplinary database of 43 million papers spanning 26 fields. For reliable experiment execution, we curate a library of 32 foundational methods for mechanism analysis. Compared with Claude Code and existing AI-scientist systems, Mechanist generates higher-quality mechanism hypotheses and executes experiments more reliably. Across four case studies, Mechanist autonomously discovers new model behaviors and their underlying mechanisms, and translates these discoveries into mechanism-guided interventions and interdisciplinary design. Specifically, Mechanist first uncovers a counterintuitive safety risk in scientific laboratories, showing that unsafe traits can transfer to fine-tuned student models through apparently safe training data and emerge across modalities. Mechanist then develops a mechanism theory of belief, revealing how models represent world knowledge, form beliefs, infer the beliefs of others, and how these mechanisms emerge during pretraining. Building on this theory, Mechanist develops targeted interventions that improve model performance across diverse scenarios. Finally, Mechanist can also advance interdisciplinary discovery through mechanistic design, providing an alternative to the computationally intensive generate-and-rerank paradigm.

补充信息

↑