arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SciHazard:一种通过分解危害评分来衡量科学安全风险的基准

SciHazard: A Benchmark for Measuring Scientific Safety Risks with Decomposed Harm Scoring

Chunxiao Li, Yuan Xiong, Lijun Li, Tianyi Du, Wenlong Zhang, Lei Bai, Jing Shao

arXiv 2607.18665首次发表:更新:

发表机构

Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对大型语言模型可能传播有害科学知识的问题,引入SciHazard基准及评估框架,含多学科问题。开发分解评估程序计算DeHarm-Score,经专家验证和模型测试,显示该方法能有效衡量风险,还发现深度研究代理存在安全盲点。

AI 中文摘要

大型语言模型(LLMs)在支持科学研究的同时,也可能将有害科学知识转化为可操作的误用指南。现有基准往往依赖与现实世界危害脱节的模板化查询,且采用缺乏领域基础的LLM作为评判范式。为解决此问题,我们引入了SciHazard,一个基于现实世界的科学风险基准和用于衡量有害性的数据集无关评估框架。SciHazard包含12个学科的2400个危险问题和600个过度安全问题,查询均基于受监管实体和记录的失败场景。为计算DeHarm-Score,我们开发了一种分解评估程序,结合查询危害严重性、拒绝行为和响应级风险。对于未被拒绝的响应,进一步将响应级危害分解为可执行性(通过带重要性加权的动态清单量化)和新出现风险(通过检索增强的声明提取和综合障碍验证评估)。专家验证研究表明,DeHarm-Score比最强基线与专家注释的一致性提高了90.17%。我们在广泛的科学安全评估中对31个前沿LLMs和深度研究代理进行了基准测试。值得注意的是,深度研究代理的平均DeHarm-Score比标准LLMs高32.3%,揭示了自主代理是当前安全防御中的关键盲点。代码和数据集可在指定网址获取。

英文摘要

Large language models (LLMs) increasingly support science, but they can also convert hazardous scientific knowledge into actionable misuse guidance. Existing benchmarks often rely on templated queries disconnected from real-world hazards, and employ LLM-as-a-Judge paradigms without domain grounding. To address this, we introduce SciHazard, a real-world-grounded benchmark for scientific risks and a dataset agnostic evaluation framework for measuring harmfulness. SciHazard contains 2400 hazardous questions and 600 oversafety questions across 12 disciplines, with both queries grounded in regulated entities and documented failure scenarios. To compute \textsc{DeHarm-Score} , we develop a decomposed evaluating procedure that combines query hazard severity, refusal behavior, and response-level risk. For non-refused responses, it further decomposes response-level harm into \textsc{Executability}, quantified via dynamic checklists with importance weighting, and \textsc{Net-new risk}, assessed through retrieval-augmented claim extraction and synthesis-barrier verification. An expert-validation study shows that \textsc{DeHarm-Score} improves agreement with expert annotations by 90.17\% over the strongest baseline. We benchmark 31 frontier LLMs and deep research agents in an extensive scientific safety evaluation. Notably, deep research agents yield 32.3\% higher mean \textsc{DeHarm-Score} than standard LLMs, exposing autonomous agents as a critical blind spot in current safety defenses. Code and dataset are available at https://anonymous.4open.science/r/DeharmScore-7B55.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑