发表机构
Capital One(第一资本)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
FinRT提出一种结构化框架,将自适应红队策略蒸馏为可复用对抗生成器,在消费金融领域显著提升攻击成功率与严重性,并保持语义多样性。
AI 中文摘要
在消费金融等受监管行业中,看似无害的用户查询可能利用大型语言模型的漏洞,引发安全故障,并将响应推向接近政策限制的危险边缘。现有的自动化红队方法在攻击有效性与生成成本之间进行权衡,同时将覆盖率、严重性和多样性视为附带目标而非联合目标。我们引入了FinRT,一个结构化框架,从自适应红队策略中构建可复用的对抗性提示生成器。在消费金融领域的六个受害者模型上,FinRT显著优于自适应搜索基线,同时将面向目标的攻击生成摊销为可复用的生成器。FinRT的攻击成功率几乎是自适应基线Rainbow Teaming的两倍(32.9%对17.2%),最大对抗性严重性提高了33%,并保持了与迭代搜索方法相当的策略域内语义多样性。我们的方法实现了高跨模型迁移性,同时表现出不同的受害者家族专业化模式。
英文摘要
In regulated industries like consumer finance, seemingly harmless user queries can exploit large language model vulnerabilities, triggering safety failures and pushing responses dangerously close to policy limits. Existing automated red-teaming methods trade off attack effectiveness against generation cost, while treating coverage, severity, and diversity as incidental rather than joint objectives. We introduce FinRT, a structured framework that builds reusable adversarial prompt generators from adaptive red-teaming strategies. Across the six victim models in consumer finance, FinRT substantially outperforms adaptive search baselines while amortizing target-facing attack generation into a reusable generator. FinRT nearly doubles the attack success rate over the adaptive baseline Rainbow Teaming (32.9% vs. 17.2%), increases maximum adversarial severity by 33%, and preserves comparable intra-policy-domain semantic diversity to iterative search methods. Our method achieves high cross-model transferability while exhibiting distinct victim-family specialization patterns.