arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-12-25 至 2025-12-25 共收录 3 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 越狱攻击 3 篇

2504.02080 2025-12-25 cs.CR cs.AI 83%

Evolving Security in LLMs: A Study of Jailbreak Attacks and Defenses

LLM安全性的演变:对劫持攻击及防御的研究

Zhengchun Shang, Wenlan Wei, Weiheng Bai

机构 * Cornell University Ithaca, NY Department of Computer Science \& Engineering University of Minnesota -- Twin Cities Minneapolis, MN

专题命中 越狱攻击 :jailbreak(title,abstract);safety(abstract);分类 cs.AI

AI总结 本文研究了LLM安全性的演变,分析了劫持攻击的检测技术,并探讨了模型版本、大小及多防御策略对安全性的综合影响。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21236 2025-12-25 cs.CR cs.AI cs.SE 81%

Casting a SPELL: Sentence Pairing Exploration for LLM Limitation-breaking

为LLM突破限制而探索句子配对:恶意代码生成的测试框架

Yifan Huang, Xiaojun Jia, Wenbo Guo, Yuqiang Sun, Yihao Huang, Chong Wang, Yang Liu

机构 * Nanyang Technological University(南洋理工大学) National University of Singapore(国立新加坡大学)

专题命中 越狱攻击 :alignment(abstract);safety(abstract);jailbreak(abstract);AI safety(abstract)

AI总结 SPELL框架通过智能句子配对探索LLM在恶意代码生成中的安全漏洞,有效提升对抗测试效果。

Comments Accepted to FSE 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.01020 2025-12-25 cs.CR cs.LG 70%

AutoAdv: Automated Adversarial Prompting for Multi-Turn Jailbreaking of Large Language Models

AutoAdv:多轮对抗性提示生成用于大型语言模型的多轮 Jailbreaking 攻击

Aashray Reddy, Andrew Zagula, Nicholas Saban

专题命中 越狱攻击 :safety(abstract);jailbreak(abstract);分类 cs.LG

AI总结 AutoAdv 提出了一种自动化多轮对抗性提示生成方法,通过策略性重写和优化配置,实现对大型语言模型的安全机制的高效攻击,揭示了其在有害内容生成上的高成功率。

Comments We encountered issues with the paper being hosted under my personal account, so we republished it under a different account associated with a university email, which makes updates and management easier. As a result, this version is a duplicate of arXiv:2511.02376

详情

展开后加载摘要…

URL PDF HTML 收藏