arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用大语言模型智能体对概念擦除进行压力测试

Stress Testing Concept Erasure with Large Language Model Agents

Yuyang Xue, Feng Chen, Zhihua Liu, Edward Moroshko, Jingyu Sun, Steven McDonagh, Sotirios A. Tsaftaris

arXiv 2607.17890首次发表:更新:

发表机构

School of Engineering, University of Edinburgh; The University of Manchester; The University of Melbourne(爱丁堡大学工程学院; 曼彻斯特大学; 墨尔本大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对概念擦除评估面临的验证挑战,提出STACE框架,利用大语言模型智能体迭代生成并验证压力测试假设,通过一套指标评估性能效率,经实验证明该框架在多方面表现优异且可拓展到其他领域。

AI 中文摘要

概念擦除旨在从训练好的生成模型中去除语义概念,这对负责任的人工智能部署愈发重要。然而,验证模型是否已稳健地去除目标概念仍是一项关键挑战。现有评估方法通常是预定义且静态的,无法揭示在各种自然语言探测和具有挑战性条件下的漏洞。此外,手动设计的评估策略可能存在偏差且难以扩展。我们认为概念擦除评估最好被表述为一种自适应假设搜索,由智能体迭代地提出、批判和验证测试以系统地扩大故障模式的覆盖范围来实施。为此,我们提出了用于概念擦除的压力测试智能体(STACE)框架,它通过基于外部知识迭代地生成和验证压力测试假设,使用多个大语言模型智能体自主地对概念擦除模型进行压力测试。我们还引入了一套用于评估基于大语言模型智能体的压力测试框架的性能和效率的指标。我们的广泛实验表明,STACE在四个概念类别上优于五个基于大语言模型的评估基线。对两个文本到图像模型、六种概念擦除方法和各种擦除强度的进一步分析表明,STACE在不同设置下都很稳健。我们还表明,STACE可以超越概念擦除评估应用于其他问题领域,如大语言模型越狱。我们的代码可匿名获取。

英文摘要

Concept erasure aims to remove semantic concepts from a trained generative model and is increasingly important for responsible AI deployment. However, verifying whether a model has robustly removed targeted concepts remains a critical challenge. Existing evaluation methods are typically pre-defined and static, failing to expose vulnerabilities under diverse natural-language probes and challenging conditions. Moreover, manually designed evaluation strategies can be biased and difficult to scale. We posit that concept erasure evaluation is best formulated as an adaptive hypothesis search, operationalised by agents that iteratively propose, critique, and verify tests to systematically expand coverage of failure modes. To this end, we propose Stress Testing Agents for Concept Erasure (STACE), a framework that autonomously stress-tests concept-erased models using multiple Large Language Model (LLM) agents, by iteratively generating and verifying stress-testing hypotheses grounded by external knowledge. We also introduce a suite of metrics for assessing the performance and efficiency of LLM-agent-powered stress-testing frameworks. Our extensive experiments show that STACE outperforms five LLM-based evaluation baselines on four concept categories. Further analysis across two T2I models, six concept erasure approaches, and various erasure strengths show that STACE is robust for different settings. We also show that STACE can be adapted beyond concept erasure evaluation to other problem domains, such as LLM jailbreaking. Our code is available anonymously.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑