arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

智能体自动驾驶显微镜基准测试支持性能验证,但不一定能泛化到未见任务

Agentic self-driving microscopy benchmarks support qualification but do not necessarily generalize to unseen tasks

Nathan S Johnson, Ian Abshire

arXiv 2608.05266首次发表:更新:

发表机构

Carl Zeiss Research Microscopy Solutions(卡尔蔡司研究显微镜解决方案)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究开发基准框架评估显微镜智能体架构等因素的性能,发现其虽可用于验证比较,但现有基准无法预测智能体在未见显微镜任务上的表现。

AI 中文摘要

大型语言模型智能体正被越来越多地开发用于控制各类科学表征工具,包括显微镜和同步加速器光束线。物理基础设施的智能体控制研究尚处于初期阶段,针对如何构建智能体系统,目前缺乏成熟的范式。设计显微镜智能体时需做出诸多选择,包括大型语言模型(LLM)的选择、使用的智能体数量、智能体职责与委派规则、检索增强生成(RAG)参数等。在设计和优化智能体显微镜控制器时,研究人员不仅希望确保智能体能正确执行已知任务,还希望其能泛化到未遇到的新任务。本研究开发了一套基准测试与轨迹记录框架,该框架可揭示:a)不同智能体架构选择如何影响显微镜任务的性能;b)基准测试在预测特定智能体是否能在未见显微镜任务中表现良好方面存在的局限性。该框架被用于评估1智能体、2智能体和3智能体图拓扑结构、5种LLM、RAG与上下文参数,以及53项显微镜基准测试中的操作约束。总共记录了105种智能体配置、1949次单独测试运行和49109次RAG检索。直接比较显示,不同配置在延迟、令牌使用量、成本和失败模式方面存在明显差异。然而,基于智能体架构和测试结果训练的代理模型无法可靠预测智能体在新的未见任务上的性能。这些结果表明,这些基准测试可用于性能验证、回归测试、诊断和直接比较,但当前异构测试套件不支持任务无关的全局配置模型。

英文摘要

Large language model agents are increasingly being developed to control a wide range of scientific characterization tools including microscopes and synchrotron beamlines. Research into agentic control of physical infrastructure is nascent and there are few well-established paradigms for how to engineer an agentic system. There are many choices to make when designing a microscopy agent, including the choice of LLM, the number of agents to use, agent responsibilities and delegation rules, retrieval-augmented generation parameters, and more. When designing and optimizing an agentic microscope controller, researchers not only want to ensure that the agent can correctly perform known tasks but also that the agent can generalize to new tasks that it has not encountered before. In this study, we develop a benchmark and trace-logging framework that reveals a) how different choices of agent architecture impact performance at microscopy tasks and b) the limitations of benchmarks for predicting if a particular agent will perform well on unseen microscopy tasks. The framework was used to evaluate one-, two-, and three-agent graph topologies, five LLMs, RAG and context parameters, and operational constraints across 53 microscopy benchmark tests. In total, 105 agent configurations, 1,949 individual test runs, and 49,109 RAG retrievals were recorded. Direct comparisons showed clear differences in latency, token use, cost, and failure mode between configurations. However, surrogate models trained on agent architecture and test results did not reliably predict an agent's performance on new, unseen tasks. These results show that these benchmarks are useful for qualification, regression testing, diagnosis, and direct comparison, but the current heterogeneous test suite does not support a task-independent global configuration model.

Comments20 pages, 8 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑