AI 中文总结
该基准评估化学与材料智能体在长时程攻击下的安全性,发现现有防御无法全面防止危险方案发布,揭示了安全评估的迫切需求。
AI 中文摘要
化学与材料智能体将文献检索、候选生成、性质预测和方案规划整合到连续发现工作流中。因此,相关的安全问题正从模型是否回答危险问题转变为智能体是否通过工具介导的工作流发布危险方案。我们引入了\ench,一个评估化学与材料智能体是否可通过用户输入、工具观察或持久记忆被引导至危险终点的基准。该基准包含432个固定的有害案例规格,涵盖八类危害、三种场景外壳、四种工具与记忆环境、一个单轮直接攻击基线,以及五种在线长时程攻击:意图劫持、工具链式利用、目标漂移、任务注入和记忆投毒。每种在线攻击的具体语言在运行时根据不断演化的轨迹生成,因此不计入静态基准规模。在固定攻击者的四模型主实验中,智能体在25.6%的运行中发布了完整的危险合成或制备程序。替换攻击者模型后,平均成功率在18.4%至26.5%之间,表明该风险并非单一攻击者的产物。从通用智能体安全改编的输入级和状态级防御,以及为化学与材料设计的候选检查,减少了一些失败,但完整路径发布率仍介于9.2%和22.5%之间。因此,现有防御无法同时覆盖多入口污染、工具状态和最终产物边界。这些结果凸显了科学智能体的快速发展与化学与材料社区可用的安全评估和防御之间的差距日益扩大。
英文摘要
Chemistry and materials agents integrate literature retrieval, candidate generation, property prediction, and protocol planning into continuous discovery workflows. Consequently, the relevant safety question is shifting from whether a model answers a hazardous question to whether an agent releases a hazardous protocol through a tool-mediated workflow. We introduce \bench, a benchmark that evaluates whether chemistry and materials agents can be steered toward hazardous endpoints through user input, tool observations, or persistent memory. The benchmark contains 432 fixed harmful case specifications spanning eight hazard classes, three scenario shells, four tool-and-memory environments, a single-turn direct-attack baseline, and five online long-horizon attacks: intent hijacking, tool chaining, objective drifting, task injection, and memory poisoning. The concrete language of each online attack is generated from the evolving trajectory at runtime and is therefore not counted in the static benchmark size. In the four-model main experiment with a fixed attacker, agents release complete hazardous synthesis or preparation procedures in 25.6\% of runs. Replacing the attacker model yields mean success rates from 18.4\% to 26.5\%, indicating that the risk is not an artifact of a single attacker. Input- and state-level defenses adapted from general-purpose agent safety, as well as candidate checks designed for chemistry and materials, reduce some failures but still leave complete-path release rates between 9.2\% and 22.5\%. Existing defenses therefore do not simultaneously cover multi-entry contamination, tool state, and the final artifact boundary. These results highlight a widening gap between the rapid development of scientific agents and the safety evaluation and defenses available to the chemistry and materials community.