Alignment is Localized: A Causal Probe into Preference Layers
机构 * Independent(独立研究者)
专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);safety(abstract);分类 cs.CL、cs.LG
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
机构 * Independent(独立研究者)
专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);safety(abstract);分类 cs.CL、cs.LG
机构 * The Alan Turing Institute(艾伦·图灵研究所) ; Statistical Laboratory(统计实验室) ; University of Cambridge(剑桥大学) ; Massachusetts Institute of Technology(麻省理工学院) ; IBM Research USA(IBM美国研究)
专题命中 偏好对齐 :RLHF(title,abstract);alignment(abstract);分类 cs.LG
Comments Extra experiments are added in new version
机构 * Graduate School of AI, KAIST(人工智能研究生院,韩国科学技术院)
专题命中 偏好对齐 :alignment(title,abstract);分类 cs.AI、cs.LG
Comments ICCV 2025 accepted
机构 * University of Science and Technology of China(中国科学技术大学) ; State Key Laboratory of Communication Content Cognition(通信内容认知国家重点实验室)
专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL
Comments NeurIPS 2025
机构 * Stanford University(斯坦福大学)
专题命中 偏好对齐 :alignment(abstract);RLHF(abstract);DPO(abstract);分类 cs.LG
机构 * College of Computer Science and Technology, Dalian University of Technology(大连理工大学计算机科学与技术学院) ; Air Force Communications NCO Academy(空军通信NCO学院)
专题命中 偏好对齐 :alignment(title);分类 cs.CL、cs.AI
Comments Accepted as a short paper at BlBM2025
专题命中 偏好对齐 :alignment(abstract);DPO(abstract);分类 cs.CL、cs.AI、cs.LG
Comments NeurIPS 2025
机构 * Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳)) ; University of Macau(澳门大学) ; University of Surrey(Surrey大学) ; Nanjing University(南京大学) ; Nanyang Technological University(南洋理工大学) ; National University of Singapore(新加坡国立大学)
专题命中 偏好对齐 :alignment(abstract);DPO(abstract);分类 cs.CL、cs.AI
Comments Accepted by NeurIPS 2025; Project Page: https://walkermitty.github.io/VimoRAG
机构 * University of Michigan - Ann Arbor(密歇根大学安阿伯分校) ; University of Toronto(多伦多大学) ; Vector Institute(向量研究所) ; MPI for Intelligent Systems, Tubingen, Germany(图宾根德国智能系统研究所)
专题命中 偏好对齐 :alignment(abstract);分类 cs.CL
Comments Preprint (Paper under review)
机构 * Rensselaer Polytechnic Institute(拉特兰理工学院)
专题命中 偏好对齐 :alignment(abstract);分类 cs.LG
机构 * School of Computing and Augmented Intelligence(计算与增强智能学院) ; Arizona State University(亚利桑那州立大学)
专题命中 偏好对齐 :DPO(abstract)
专题命中 安全训练 :alignment(abstract);trustworthy(abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 安全训练 :safety(abstract);trustworthy(abstract);分类 cs.AI
机构 * EPFL(苏黎世联邦理工学院) ; LatentWorlds AI ; TUDelft(代尔夫特理工大学) ; Embodied AI SA(具身人工智能股份有限公司)
专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG
Comments Accepted by NeurIPS 2025 SpaVLE workshop. 4 pages, 2 figures(in main paper, excluding references and supplements)
专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.CY
机构 * MATS & Anthropic Fellows Program(MATS与Anthropic Fellow项目) ; Thinking Machines Lab(Thinking Machines实验室) ; Anthropic
专题命中 安全训练 :safety(abstract);分类 cs.AI
专题命中 安全训练 :safety(abstract)
Comments 52 pages; 3 figures; PRIMEarxiv template; fully reproducible artifact (code, configs, plots)
机构 * University of Bonn(波恩大学) ; Lamarr Institute for Machine Learning and Artificial Intelligence(拉马尔人工智能与机器学习研究所) ; Center for Robotics(机器人中心) ; University of Mainz(美因茨大学) ; Fraunhofer Institute for Intelligent Analysis and Information Systems IAIS(弗劳恩霍夫智能分析与信息系统研究所)
专题命中 安全训练 :trustworthy(abstract)
机构 * Humanoid Robots Lab, University of Bonn(波恩大学人形机器人实验室)
专题命中 安全训练 :safety(abstract)
专题命中 越狱攻击 :jailbreak(title,abstract);alignment(abstract);safety(abstract);分类 cs.AI
Comments Withdrawn due to an accidental duplicate submission. This paper (arXiv:2510.15430) was unintentionally submitted as a new entry instead of a new version of our previous work (arXiv:2508.09201)
机构 * MBZUAI(马克斯·普朗克人工智能研究所) ; University of Edinburgh(爱丁堡大学)
专题命中 越狱攻击 :jailbreak(title,abstract);safety(abstract);分类 cs.CL
机构 * MBZUAI Abu Dhabi, UAE(阿布扎赫德MBZUAI)
专题命中 越狱攻击 :safety(abstract);jailbreak(abstract);分类 cs.CL、cs.LG
Comments NeurIPS 2025 (spotlight)
机构 * Department of Electrical and Electronic Engineering, The Hong Kong Polytechnic University(香港理工大学电子与电气工程系) ; School of Electronics and Information, Northwestern Polytechnical University(西北工业大学电子与信息学院)
专题命中 越狱攻击 :alignment(abstract);分类 cs.CL、cs.AI
专题命中 越狱攻击 :safety(abstract);分类 cs.AI、cs.CY
Comments Preprint V3 (October 2025)
机构 * ETH Zurich(苏黎世联邦理工学院) ; University at Buffalo(布法罗大学) ; Bytedance, Security Research(字节跳动安全研究)
专题命中 越狱攻击 :prompt injection(abstract);分类 cs.LG
专题命中 提示注入 :prompt injection(title,abstract);分类 cs.CY
Comments Accepted as a poster to Soups 2025
Journal ref The Twenty-First Symposium on Usable Privacy and Security (SOUPS 2025) Poster
机构 * University at Buffalo(布法罗大学)
专题命中 提示注入 :jailbreak(abstract);prompt injection(abstract);分类 cs.AI
专题命中 提示注入 :prompt injection(abstract)
Journal ref PACMI'2025
机构 * Department of Electrical and Computer Engineering & Ingenuity Labs Research Institute(电气与计算机工程系及创新实验室研究机构) ; Conflict Analytics Lab, Queen’s University(冲突分析实验室,女王大学) ; Cornell Law School(康奈尔法学院)
专题命中 幻觉与事实性 :alignment(title,abstract);分类 cs.CL
Comments Findings of EMNLP 2025, 5 pages
机构 * Wroclaw University of Science and Technology(沃拉日-克拉夫大学科学与技术学院) ; University of Technology Sydney(悉尼技术大学)
专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL、cs.AI、cs.LG
Comments Accepted to EMNLP 2025. Code available at https://github.com/graphml-lab-pwr/lapeigvals