arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-07-29 至 2025-07-29 共收录 4 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 越狱攻击 4 篇

2507.20333 2025-07-29 cs.AI cs.LG stat.ML 88%

The Blessing and Curse of Dimensionality in Safety Alignment

Rachel S. Y. Teo, Laziz U. Abdullaev, Tan M. Nguyen

机构 * Department of Mathematics National University of Singapore(数学系新加坡国立大学)

专题命中 越狱攻击 :alignment(title,abstract);safety(title,abstract);分类 cs.AI、cs.LG

Comments Published as a conference paper at COLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15606 2025-07-29 cs.LG cs.AI cs.CL 85%

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning

Gabriel J. Perin, Runjin Chen, Xuxi Chen, Nina S. T. Hirata, Zhangyang Wang, Junyuan Hong

机构 * University of São Paulo(圣保罗大学) University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 越狱攻击 :safety(title,abstract);alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19880 2025-07-29 cs.CR cs.AI 57%

Trivial Trojans: How Minimal MCP Servers Enable Cross-Tool Exfiltration of Sensitive Data

Nicola Croce, Tobin South

机构 * Pivotal Research(Pivotal研究机构) Stanford University(斯坦福大学)

专题命中 越狱攻击 :trustworthy(abstract);分类 cs.AI

Comments Abstract submitted to the Technical AI Governance Forum 2025 (https://www.techgov.ai/)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19609 2025-07-29 cs.CR 50%

Securing the Internet of Medical Things (IoMT): Real-World Attack Taxonomy and Practical Security Measures

Suman Deb, Emil Lupu, Emm Mic Drakakis, Anil Anthony Bharath, Zhen Kit Leung, Guang Rui Ma, Anupam Chattopadhyay

专题命中 越狱攻击 :safety(abstract)

Comments Submitted as a book chapter in 'Handbook of Industrial Internet of Things' to be published by Springer Nature

详情

展开后加载摘要…

URL PDF HTML 收藏