研究中的AI智能体:智能体系统如何改变科学工作
Research with AI Agents: How Agentic Systems Are Changing Scientific Work
- Fraunhofer Institute for Digital Medicine MEVIS(弗劳恩霍夫数字医学MEVIS研究所)
- Constructor University(康斯特鲁克特大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文探讨智能体AI系统在科学研究中的应用条件与效果,指出其将工作从执行转向指导与审查,并强调研究人员对结果的责任及专家判断在评估研究问题重要性中的不可替代性。
AI中文摘要:
背景。智能体AI系统自主地将文献检索、数据分析和编程等任务分解为子任务,搜索网络、访问数据库并执行代码。这使得它们能够高速执行数字研究任务。目标。在什么条件下使用智能体系统能产生可靠的效率提升,哪些任务仍由研究人员完成?材料与方法。对当前关于文献检索、数据分析、软件开发及临床决策支持的研究进行总结。结果。对于数字活动,工作从执行转向指导和审查。当预期行为可以事先形式化并自动测试时,效率提升最大。在复杂的智能体系统中,记录下来的推理步骤和工具调用序列可能迅速变得过于庞大,难以进行人工审查。此外,模型生成的解释并不能可靠地反映输出是如何产生的。向更可靠系统迈进的一个可能步骤是对各个组件进行验证。有限的审查能力不仅限于研究本身;科学文章和拨款提案的审查也正达到能力极限。在病理学领域,经过整理和标注的数据、研究人员自身的分析技能以及机构间的经验交流正变得越来越重要。结论。研究人员仍需对其结果负责。他们必须决定将哪些任务委托出去以及如何审查结果。研究问题对患者、领域和社会的重要性无法完全通过形式化标准来评估,这仍属于专家判断的范畴。智能体系统可以为此腾出时间。
英文摘要:
Background. Agentic AI systems independently decompose tasks such as literature search, data analysis, and programming into subtasks, search the web, access databases, and execute code. This allows them to perform digital research tasks at high speed. Objectives. Under what conditions does the use of agentic systems produce reliable efficiency gains, and which tasks remain with researchers? Materials and methods. Summary of current studies on literature searches, data analysis, software development, and clinical decision support. Results. For digital activities, work shifts from execution to steering and review. Efficiency gains are greatest when expected behavior can be formalized in advance and tested automatically. In complex agentic systems, recorded sequences of reasoning steps and tool calls can quickly become too extensive for human review. Furthermore, explanations generated by the model do not reliably reflect how an output was produced. One possible step toward more reliable systems is the validation of individual components. The limited reviewability extends beyond research itself; the review of scientific articles and grant proposals is also reaching capacity limits. In pathology, curated and annotated data, researchers' own analytical skills, and institutional exchange of experience are becoming increasingly important. Conclusions. Researchers remain responsible for their results. They must determine what to delegate and how to review the results. The importance of a research question to patients, the field, and society cannot be fully assessed using formalized criteria and remains a matter of expert judgment. Agentic systems can free up time for this.