arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TIER:用于评估大语言模型安全行为的威胁隐式基准

TIER: Threat Implicitness Benchmark for Evaluating LLM Safety Behaviors

Thu-Hien Trinh-Thi, Hai-Yen Vong, Thanh-Ha Ung-Dung, Tram Ho

arXiv 2609.05117首次发表:更新:

发表机构

Faculty of Information Technology, University of Science, Vietnam National University, Ho Chi Minh City(胡志明市越南国家大学理学院信息技术系)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对现有LLM安全基准忽略威胁隐式程度的问题,提出TIER基准,经6个开源LLM实验发现安全行为随威胁级别渐变,且相似攻击成功率模型响应分布不同,凸显行为感知评估的必要。

AI 中文摘要

当前大语言模型(LLM)安全基准大多依赖二元指标,忽略了模型对具有不同威胁隐式程度的有害提示的响应方式。我们推出TIER(Threat Implicitness Benchmark,威胁隐式基准),用于评估LLM的行为安全。TIER涵盖四个风险领域和四个威胁级别,从明确的有害请求到复杂的越狱提示。响应采用六标签行为量表和两名独立LLM评判进行评估。对六个开源权重LLM的实验表明,安全行为随威胁级别逐步演变,而非直接从拒绝转向服从。上下文提示产生最多样化的行为,而越狱提示则暴露出最大的鲁棒性差距。此外,具有相似攻击成功率的模型可能表现出不同的响应分布,凸显了基于行为的LLM安全评估的必要性。

英文摘要

Current LLM safety benchmarks largely rely on binary metrics, overlooking how models respond to harmful prompts with varying threat implicitness. We introduce TIER, a Threat Implicitness Benchmark for behavioral safety evaluation of LLMs. TIER covers four risk domains and four threat levels, from explicit harmful requests to sophisticated jailbreaks. Responses are assessed using a six-label behavior scale and two independent LLM judges. Experiments on six open-weight LLMs show that safety behaviors evolve gradually across threat levels rather than shifting directly from refusal to compliance. Contextual prompts yield the most diverse behaviors, while jailbreaks reveal the largest robustness gaps. Furthermore, models with similar Attack Success Rates can exhibit distinct response distributions, highlighting the need for behavior-aware LLM safety evaluation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑