arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

BioDisclose:对抗性诱导下生物医学安全的可操作性感知基准

BioDisclose: An Actionability-Aware Benchmark for Biomedical Safety under Adversarial Elicitation

Yinuo Zhu, He Liu, Boyuan Gu

arXiv 2607.25700首次发表:更新:

AI 中文总结

研究针对大语言模型在对抗性诱导下生物医学知识披露行为不足问题,引入BioDisclose基准,含多领域多类别提示,对模型响应四级评分,通过实验揭示不同模型、策略及风险类别差异,为评估生物医学安全提供新测试平台。

AI 中文摘要

大语言模型(LLMs)越来越多地支持生物医学研究,但其在对抗性获取两用知识请求下的行为仍未得到充分描述。我们引入了BioDisclose,这是一个用于衡量对抗性诱导下生物医学知识披露的基准。BioDisclose包含480个提示,来自六个生物医学风险领域的24个专家编写的场景以及学术、历史、角色扮演和分解提示四个诱导类别。我们从拒绝到可执行披露的四级量表对模型响应进行评分,区分高级讨论与技术上具体且可操作的内容,包括拒绝后泄露行为。在五个已部署的LLM系统中,详细或更高披露率差异很大,从9.2%到64.0%不等。学术框架平均是最有效的诱导类别(43.2%),而实验室安全场景在各领域中披露率最高(51.5%)。这些结果揭示了模型、提示策略和生物医学风险类别之间的显著差异,表明当前在高风险科学环境中的保障措施仍然不均衡。BioDisclose为评估生物医学安全提供了一个聚焦的测试平台,超越了二元拒绝指标。

英文摘要

Large language models (LLMs) increasingly support biomedical research, yet their behavior under adversarial requests for dual-use knowledge remains insufficiently characterized. We introduce BioDisclose, a benchmark for measuring biomedical knowledge disclosure under adversarial elicitation. BioDisclose contains 480 prompts derived from 24 expert-authored scenarios across six biomedical risk domains and four elicitation families spanning academic, historical, role-playing, and decomposed prompting. We grade model responses on a four-level scale from refusal to executable disclosure, distinguishing high-level discussion from technically specific and actionable content, including refuse-then-leak behavior. Across five deployed LLM systems, detailed-or-higher disclosure rates vary substantially, ranging from 9.2% to 64.0%. Academic framing is the most effective elicitation family on average (43.2%), while laboratory safety scenarios show the highest disclosure rate across domains (51.5%). These results reveal pronounced variation across models, prompting strategies, and biomedical risk categories, suggesting that current safeguards remain uneven in high-stakes scientific settings. BioDisclose provides a focused testbed for evaluating biomedical safety beyond binary refusal metrics.

Comments27 pages, 3 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑