TombRaider: Entering the Vault of History to Jailbreak Large Language Models
机构 * UNSW(新南威尔士大学) ; NTU(国立大学)
专题命中 越狱攻击 :jailbreak(title,abstract);safety(abstract);red teaming(abstract);分类 cs.CL、cs.AI、cs.CY
Comments Main Conference of EMNLP
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
机构 * UNSW(新南威尔士大学) ; NTU(国立大学)
专题命中 越狱攻击 :jailbreak(title,abstract);safety(abstract);red teaming(abstract);分类 cs.CL、cs.AI、cs.CY
Comments Main Conference of EMNLP
机构 * Duke Kunshan University(杜克昆山大学)
专题命中 越狱攻击 :jailbreak(title,abstract);alignment(abstract);分类 cs.CL、cs.AI
Comments ICONIP 2025
机构 * Huazhong University of Science and Technology(华中科技大学) ; University of Notre Dame(圣母大学) ; Lehigh University(莱文森大学) ; Duke University(杜克大学)
专题命中 越狱攻击 :prompt injection(title,abstract);jailbreak(abstract);分类 cs.AI
Comments To appear in the Proceedings of The ACM Conference on Computer and Communications Security (CCS), 2024
机构 * University of Strathclyde(斯特拉思克莱德大学)
专题命中 越狱攻击 :alignment(abstract,comments);safety(abstract);jailbreak(abstract);分类 cs.CL、cs.AI
Comments Published at Transaction of Machine Learning Research 08/2025, Large Language Models (LLMs), Interference-time activation shifting, Steerability, Explainability, AI alignment, Interpretability
专题命中 越狱攻击 :prompt injection(title,abstract)
机构 * Nanyang Technological University(南洋理工大学) ; Institute of High Performance Computing(高性能计算研究所) ; Agency for Science, Technology and Research(科技研究局)
专题命中 越狱攻击 :safety(abstract);jailbreak(abstract);分类 cs.CL
专题命中 越狱攻击 :safety(abstract);分类 cs.AI
Comments This is the authors accepted manuscript of an article accepted for publication in Cluster Computing. The final published version is available at: 10.1007/s10586-025-05326-9