arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.00578cs.AIcs.CR

相同请求,不同边界:评估对话语境下的网络安全辅助能力

Same Request, Different Boundary: Evaluating Cybersecurity Assistance across Conversational Contexts

Rui Yang, Yang Hong, Yichao Xu, Zhengyu Liu, Ziyang Li, Yinzhi Cao

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出3R-Bench基准数据集,评估8个LLMs在不同对话语境下对网络安全请求的合规性,发现对话语境会显著改变模型对同一请求的响应,失败反馈仅能小幅恢复合规性损失。

中文摘要 AI 辅助

大型语言模型(LLMs)可解决复杂问题,但其在高风险领域的滥用会导致严重后果,因此模型提供方会限制对潜在有害请求的辅助。若拒绝所有网络安全请求,会损害合法用户的利益,提供方需要一种机制来阻止恶意使用,同时不拒绝防御者的合法辅助。现有的网络安全专用数据集可评估该机制,但均未考虑请求的对话语境。本文引入3R-Bench(Refusal, Repetition, and Revision),这是一个包含150个真实网络安全请求的基准数据集,新增两种对抗性对话设置,并在该基准上评估8个LLMs。未改变的请求会因之前的辅助行为而产生截然不同的响应:在400对面板数据中,有376对可用,合规率从被拒绝历史后的62.0%上升至被接受历史后的85.1%;在对话分解下则呈现相反模式,直接响应的合规率为501/800,经过对话后降至172/800,其中738对在两种条件下均返回模型生成文本,降幅达45.1个百分点;失败反馈仅能恢复一小部分损失。

英文摘要

Large Language Models (LLMs) can solve complex problems, but their misuse in high-risk domains can lead to severe consequences. Model providers therefore restrict assistance for potentially harmful requests. Refusing all cybersecurity requests would therefore harm legitimate users. Providers need a mechanism to block malicious use without denying legitimate assistance to defenders. Existing cybersecurity-specific datasets evaluate this mechanism, but none considers the conversational context of a request. We introduce 3R-Bench (Refusal, Repetition, and Revision), a benchmark of 150 real-world cybersecurity requests augmented with two adversarial conversational settings, and evaluate eight LLMs on it. Prior assistant behavior strongly changes responses to an unchanged request: among 376 available pairs from a 400-pair panel, compliance rises from 62.0% after refused history to 85.1% after accepted history. The opposite pattern appears under dialogue decomposition. In comparison, compliance falls from 501/800 direct responses to 172/800 after dialogue; among 738 pairs returning model-authored text in both conditions, the decrease is 45.1 points. Failure feedback recovers only a small fraction of this loss.

发表机构

  • Johns Hopkins University(约翰斯·霍普金斯大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑