发表机构
Stanford University; Tri-institutional Center for Translational Research in Neuroimaging and Data Science (TReNDS); Georgia State University; Georgia Institute of Technology; Emory University; Inria; CEA; Université(斯坦福大学; 三机构神经成像与数据科学转化研究中心(TReNDS); 佐治亚州立大学; 佐治亚理工学院; 埃默里大学; 法国国家信息与自动化研究所; 法国原子能和替代能源委员会; 大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究推出Brain Researcher智能体研究平台,用于神经影像数据分析,其遵循分析规则等要求,提升了工具选择准确率与可验证依据,嵌入方法判断以增强智能体AI的分析严谨性。
AI 中文摘要
AI智能体可执行科学分析,但仅在权衡多种替代方案并将主张限定于证据支持的范围内时,其分析输出才能成为可辩护的论断。智能体可能会重现包括选择性分析、过早宣告成功以及针对不完美标准进行优化在内的失败。我们推出Brain Researcher,这是一款在神经影像研究者的计算环境中运行的智能体研究工具,遵循可接受分析的规则、必要检查和主张范围的要求。在基准测试中,Brain Researcher使7种模型的首选工具选择准确率提高了70.2个百分点(未使用时为23.3%,使用后为93.6%),可验证的依据从4.6%提升至22.0%。在合作者主导和自我演进的研究中,多重宇宙分析揭示了分析选择的敏感性,科学审查将主张分为接受、有条件接受、修订、阻止、拒绝或推迟。通过将决策与证据和来源关联,Brain Researcher将方法判断嵌入工作流程之中,而非置于其后。
英文摘要
AI agents can execute scientific analyses, but an analytic output becomes a defensible claim only after alternatives are weighed and the claim is limited to what the evidence supports. Agents may reproduce failures including selective analysis, premature declarations of success and optimization of imperfect criteria. We present Brain Researcher, an agentic research harness operating in a neuroimaging researcher's computational environment under rules for admissible analyses, required checks and claim scope. In benchmarks, Brain Researcher increased first-choice tool-selection accuracy across seven models by 70.2 percentage points (23.3% without it versus 93.6% with it) and verifiable grounding from 4.6% to 22.0%. In collaborator-led and self-evolving studies, multiverse analyses exposed analytic-choice sensitivity, and scientific review classified claims as accepted, qualified, revised, blocked, rejected or deferred. By linking decisions to evidence and provenance, Brain Researcher embeds methodological judgment within the workflow, not after it.
Comments103 pages, 19 figures; Supplementary Information included