arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34537cs.AI

科学推理的马拉松:科学智能体在多轮交互中对扰动的鲁棒性

The Marathon of Scientific Reasoning: Robustness of Scientific Agents to Perturbations in Multi-Turn Interactions

Xiaoting Lyu, Xinbo Ma, Yufei Han, Hangwei Qian, Ziyang Lin, Bin Wang, Bin Wang, Wei Wang

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出SciARP基准,通过620个科学问题和13种扰动类型评估八个LLM科学智能体的鲁棒性,发现扰动效应与任务进展可解耦、强基线不一定鲁棒、扰动影响具有时滞传播性。

中文摘要 AI 辅助

基于大型语言模型(LLM)的科学智能体越来越多地被用于科学问题求解,然而它们在多轮交互过程中对不完善性的鲁棒性仍然知之甚少。我们引入了SciARP(科学智能体对扰动的鲁棒性),这是一个基准测试,用于在多轮问题求解过程中,在科学上合理的扰动下评估科学智能体。SciARP将620个科学问题转化为3至13轮相互依赖的任务,并定义了涵盖问题理解、证据处理、推理和结论形成四个方面的13种扰动类型。每个任务的干净版本和扰动版本在匹配的设置下独立执行,产生成对的实时轨迹,用于评估任务成功和过程可靠性。在来自四个模型家族的八个LLM上进行的实验揭示了三个关键的鲁棒性特征。首先,不同类别的科学扰动表现出不同的鲁棒性特征,并可能将任务进展与科学可靠性解耦:智能体即使在其信息或推理变得不可靠之后,仍可能继续推进任务。其次,更强的干净任务性能并不一定转化为更强的鲁棒性,因为具有更高干净任务准确率的模型在扰动下可能表现出更大的性能下降。第三,扰动效应表现出强烈的时间动态:它们可能在多个回合中保持潜伏状态,然后出现并通过下游依赖关系传播。总之,这些发现表明,当前的科学智能体对科学上合理的扰动仍然不够鲁棒,失败往往未被检测到、传播并抵抗恢复。

英文摘要

Large language model (LLM)-based scientific agents are increasingly used for scientific problem solving, yet their robustness to imperfections arising during multi-turn interactions remains poorly understood. We introduce \textsc{SciARP} (\textbf{Sci}entific \textbf{A}gent \textbf{R}obustness to \textbf{P}erturbations), a benchmark for evaluating scientific agents under scientifically plausible perturbations throughout multi-turn problem solving. \textsc{SciARP} transforms 620 scientific problems into interdependent tasks of 3--13 turns and defines 13 perturbation types spanning problem understanding, evidence processing, reasoning, and conclusion formation. Clean and perturbed versions of each task are independently executed under matched settings, producing paired live trajectories for evaluating both task success and process reliability. Experiments across eight LLMs from four model families reveal three key robustness characteristics. First, different classes of scientific perturbations exhibit distinct robustness profiles and can decouple task progression from scientific reliability: agents may continue advancing through the task even after their information or reasoning has become unreliable. Second, stronger clean-task performance does not necessarily translate into stronger robustness, as models with higher clean-task accuracy can exhibit larger degradation under perturbation. Third, perturbation effects exhibit strong temporal dynamics: they may remain latent for multiple turns before emerging and subsequently propagate through downstream dependencies. Together, these findings show that current scientific agents remain insufficiently robust to scientifically plausible perturbations, with failures often remaining undetected, propagating, and resisting recovery.

发表机构

  • Inria(法国国家信息与自动化研究所)
  • Agency for Science, Technology and Research (A*STAR)(新加坡科技研究局)
  • Zhejiang Key Laboratory of Artificial Intelligence of Things (AIoT) Network and Data Security(浙江省人工智能物联网网络与数据安全重点实验室)

机构由 AI 辅助整理,请以论文原文为准。

↑