When Alignment Hurts: Decoupling Representational Spaces in Multilingual Models
机构 * MBZUAI(马克斯·普朗克人工智能研究所) ; NICT, Japan(日本信息通信技术研究所) ; IIT Madras(印度理工学院Madras分校)
专题命中 其他安全 :alignment(title,abstract);分类 cs.CL
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
机构 * MBZUAI(马克斯·普朗克人工智能研究所) ; NICT, Japan(日本信息通信技术研究所) ; IIT Madras(印度理工学院Madras分校)
专题命中 其他安全 :alignment(title,abstract);分类 cs.CL
专题命中 其他安全 :alignment(abstract);safety(abstract)
机构 * NSF Center for Quantum Network(NSF量子网络中心) ; University of California, Los Angeles(加州大学洛杉矶分校)
专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI
Comments ICML 2025 Workshop on MAS
机构 * Stanford University(斯坦福大学)
专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG
Journal ref Transactions on Machine Learning Research (TMLR) 2835-8856 (2025)
专题命中 其他安全 :alignment(abstract);分类 cs.CY
Comments 25 pages, 4 pages
专题命中 其他安全 :alignment(abstract);分类 cs.LG
机构 * University of Toronto(多伦多大学)
专题命中 其他安全 :alignment(abstract);分类 cs.AI
Journal ref PACMHCI (CSCW 2025)
机构 * Hong Kong University of Science and Technology(香港理工大学) ; Cohere
专题命中 其他安全 :alignment(abstract);分类 cs.CL
专题命中 其他安全 :safety(abstract);分类 cs.LG
Comments Presented at the 2nd Reinforcement Learning Conference (RLC2025), Edmonton, Canada. To be published in the Proceedings of the Reinforcement Learning Journal 2025
专题命中 其他安全 :alignment(abstract);分类 cs.CL
专题命中 其他安全 :safety(abstract)
机构 * VRG, FEE, Czech Technical University in Prague(捷克布拉格技术大学)
专题命中 其他安全 :alignment(abstract)
Comments ICCVW 2025 accepted paper. Workshop name: "What is Next in Multimodal Foundation Models?"
机构 * South China University of Technology(南方科技大学) ; The Hong Kong Polytechnic University(香港理工大学) ; Key Laboratory of Big Data and Intelligent Robot Ministry of Education(教育部大数据与智能机器人重点实验室)
专题命中 其他安全 :safety(abstract)
Comments Accepted by ACM MM 2025
专题命中 其他安全 :alignment(abstract)