arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32918physics.opticscs.AIcs.CVcs.RO

视觉-语言智能体在光学实验室中的主动感知

Vision-Language Agents for Active Perception in Optics Laboratories

  • Massachusetts Institute of Technology(麻省理工学院)

机构由 AI 辅助整理,请以论文原文为准。

Ryan Lopez, Sachin Vaidya, Seou Choi, Serena Landers, Marin Soljačić

AI总结:

本研究以光学实验为平台,验证通用视觉-语言模型能通过视觉反馈闭环控制实验,主动获取信息并适应行为,成为科学决策智能体。

AI中文摘要:

视觉-语言模型(VLMs)正越来越多地被用于科学工作流程,但它们作为智能体直接根据视觉反馈控制实验室实验的能力仍未得到充分探索。这一能力之所以重要,是因为许多实验室任务并不天然提供密集的、预定义的数值目标:信息性信号可能是稀疏的、间歇性的或视觉上模糊的。一个更通用的实验室智能体应当能够解读视觉观察结果,采取行动获取有用的反馈,并根据这些行动的后果调整自身行为。我们以实验光学为测试平台,研究通用视觉-语言模型能否执行这种闭环科学控制。我们在三个隔离不同能力的实验系统上评估智能体:迈克尔逊干涉仪、双镜腔和四镜光学中继系统。智能体观察相机图像,直接发出执行器和测量命令,并在控制过程中保留其交互历史,而无需接收工程化的标量目标。在这些实验及匹配的模拟中,我们发现,在给定任务特定的自然语言指导时,视觉-语言模型能够估计执行器-响应关系,通过干预解决模糊观察,并在信号稀疏时主动创造信息丰富的视觉反馈。这些结果表明,预训练的多模态模型可以在实验回路中充当重要的决策智能体。我们的工作还确立了光学作为科学智能体中视觉推理和主动感知的物理接地测试平台。

英文摘要:

Vision-language models (VLMs) are increasingly being used in scientific workflows, but their ability as agents to directly control laboratory experiments from visual feedback remains underexplored. This capability is important because many laboratory tasks do not naturally provide dense, pre-defined numerical objectives: informative signals can be sparse, intermittent, or visually ambiguous. A more general laboratory agent should instead be able to interpret visual observations, take actions to acquire useful feedback, and adapt its behavior based on the consequences of those actions. We study whether general-purpose VLMs can perform this kind of closed-loop scientific control using experimental optics as a testbed. We evaluate agents on three experimental systems that isolate distinct capabilities: a Michelson interferometer, a two-mirror cavity, and a four-mirror optical relay. The agents observe camera images, directly issue actuator and measurement commands, and retain their interaction history without receiving an engineered scalar objective during control. Across these experiments and matched simulations, we find that, given task-specific natural-language guidance, VLMs can estimate actuator-response relationships, resolve ambiguous observations through intervention, and actively create informative visual feedback when signals are sparse. These results suggest that pretrained multimodal models can serve as important decision-making agents within the experimental loop. Our work also establishes optics as a physically grounded testbed for visual reasoning and active perception in scientific agents.

↑