arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-10-06 至 2025-10-06 共收录 4 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 越狱攻击 4 篇

2412.07192 2025-10-06 cs.CR cs.CL cs.LG 79%

PrisonBreak: Jailbreaking Large Language Models with at Most Twenty-Five Targeted Bit-flips

Zachary Coalson, Jeonghyun Woo, Chris S. Lin, Joyce Qu, Yu Sun, Shiyang Chen, Lishan Yang, Gururaj Saileshwar, Prashant Nair, Bo Fang, Sanghyun Hong

机构 * Oregon State University(俄勒冈州立大学) University of British Columbia(不列颠哥伦比亚大学) University of Toronto(多伦多大学) George Mason University(乔治·梅森大学) Rutgers University(罗格斯大学) University of Texas at Arlington(德克萨斯大学阿灵顿分校)

专题命中 越狱攻击 :alignment(abstract);safety(abstract);jailbreak(abstract);分类 cs.CL、cs.LG

Comments Pre-print

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01494 2025-10-06 cs.LG cs.AI 73%

Understanding Adversarial Transfer: Why Representation-Space Attacks Fail Where Data-Space Attacks Succeed

Isha Gupta, Rylan Schaeffer, Joshua Kazdan, Ken Ziyu Liu, Sanmi Koyejo

机构 * ETH Zürich(苏黎世联邦理工学院) Stanford CS(斯坦福大学计算机科学系)

专题命中 越狱攻击 :alignment(abstract);jailbreak(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21634 2025-10-06 cs.CR cs.AI cs.LG cs.NI 73%

MobiLLM: An Agentic AI Framework for Closed-Loop Threat Mitigation in 6G Open RANs

Prakhar Sharma, Haohuang Wen, Vinod Yegneswaran, Ashish Gehani, Phillip Porras, Zhiqiang Lin

机构 * SRI The Ohio State University(俄亥俄州立大学)

专题命中 越狱攻击 :safety(abstract);trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03204 2025-10-06 cs.CL 57%

FocusAgent: Simple Yet Effective Ways of Trimming the Large Context of Web Agents

Imene Kerboua, Sahar Omidi Shayegan, Megh Thakkar, Xing Han Lù, Léo Boisvert, Massimo Caccia, Jérémy Espinas, Alexandre Aussem, Véronique Eglin, Alexandre Lacoste

机构 * LIRIS - CNRS, INSA Lyon, Universite Claude Bernard Lyon 1(LIRIS - CNRS,INSA里昂,克劳德·贝尔纳大学里昂) Esker ServiceNow Research(ServiceNow研究) Mila - Quebec AI Institute(魁北克人工智能研究所) McGill University(麦吉尔大学) Polytechnique Montréal(蒙特利尔理工学院)

专题命中 越狱攻击 :prompt injection(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏